NewsBlogsTechTech EngineeringCanalsScienceResearchProgrammingAI

Feeds

    100RApple Machine LearningArXiv (cs.CV)BairBBC - Top StoriesBBC TechnologyBen EvansCisco EngineeringDaring FireballDatabricksDiamond GeezerDropbox TechEngineering at MetaGitHubHackadayHacker NewsHugging FaceIan VisitsIEEE SpectrumInk & SwitchInstagram EngineeringKDnuggetsLWN.netLyft EngineeringMacRumorsMattBitsNetflix EngineeringOllamaOMG UbuntuOnTheCutOpenAIPythonQuanta MagazineRachel by the BayRustSalesforce EngineeringShopify EngineeringSimon WillisonSpotify EngineeringThe ConversationThe Fly BlogThe TimesTorrentFreakTowards Data ScienceUber EngineeringWillowWiredZed

Hugging Face

  • Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
    8 hours ago
  • tokenizers v1: encode, decode and scaling, measured
    22 hours ago
  • Your Agent Aced the Task. Will It Do It Again?
    9/15/2026 at 16:00
  • Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
    9/10/2026 at 00:00
  • Rebuilding AUTOMATIC1111 with Gradio Workflow
    9/10/2026 at 00:00
  • NeoMME: an efficient Multimodal-native and Multilingual Encoder
    9/3/2026 at 13:13
  • Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
    9/3/2026 at 00:00
  • Give Your Coding Agents a Memory You Own
    9/3/2026 at 00:00
  • Training a coding model to paint watercolours with TRL and OpenEnv
    9/3/2026 at 00:00
  • BenchMIRT: What are LLM benchmarks actually measuring?
    9/1/2026 at 21:39
  • Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
    9/1/2026 at 00:00
  • The Open ASR Leaderboard Adds Its First Global South Language
    8/28/2026 at 00:00
  • Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
    8/26/2026 at 00:00
  • Granite 4.2 LLMs: How They're Built
    8/25/2026 at 15:14
  • Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
    8/25/2026 at 11:39
  • Wire It, Run It, Deploy It: AI Workflows in Gradio
    8/25/2026 at 00:00
  • How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
    8/21/2026 at 00:00
  • Measuring benchmark optimization in speech recognition
    8/21/2026 at 00:00
  • Up to 3.2x Faster Inference with LFM2.5-DSpark
    8/20/2026 at 16:52
  • How Much Memory Does Your Agent Actually Need?
    8/18/2026 at 18:09
  • Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
    8/18/2026 at 00:00
  • Same Cluster, 33 Points More Utilization: What Changed Was the Order
    8/17/2026 at 19:46
  • State of Open Models: Summer 2026 Observations
    8/14/2026 at 00:00
  • Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets
    8/13/2026 at 17:16
  • What We Learned by Reproducing 2,200 papers from ICML
    8/13/2026 at 00:00
  • Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
    8/12/2026 at 16:14
  • Thinking of ACE? We Can Do It with Fewer Tokens
    8/11/2026 at 13:37
  • Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
    8/10/2026 at 16:25
  • Making Knowledge Distillation Cheap Enough to Run at Scale
    8/10/2026 at 10:05
  • Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
    8/10/2026 at 00:00