Public
NeoMME: an efficient Multimodal-native and Multilingual Encoder
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Give Your Coding Agents a Memory You Own
Training a coding model to paint watercolours with TRL and OpenEnv
Real-Time Intelligence with IBM Time Series Models on Confluent
BenchMIRT: What are LLM benchmarks actually measuring?
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
The Open ASR Leaderboard Adds Its First Global South Language
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Granite 4.2 LLMs: How They're Built
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Wire It, Run It, Deploy It: AI Workflows in Gradio
Measuring benchmark optimization in speech recognition
How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
Up to 3.2x Faster Inference with LFM2.5-DSpark
How Much Memory Does Your Agent Actually Need?
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Same Cluster, 33 Points More Utilization: What Changed Was the Order
@huggingface – Blog | Pasteblog