Articles tagged: benchmarking

10 articles

AI tools

Introducing Real World VoiceEQ: Measuring the Human Quality of Voice AI

Real World VoiceEQ is a new benchmark that evaluates voice AI systems on human-likeness, emotional expressiveness, and natural prosody, pr...

Jul 16, 20267 min
AI tools

Behind the Scenes of Distributed Training and Why Your GPU Wiring Matters as Much as Your Strategy

Distributed training accelerates AI model development, but network topology and GPU interconnect often bottleneck performance. This articl...

Jul 11, 20267 min
AI tools

Time-Series LLMs, Explained with t0-alpha

Time-series LLMs like t0-alpha leverage transformer architectures to analyze sequential data. This article explains how t0-alpha handles f...

Jul 5, 20266 min
AI research

ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

ScarfBench introduces a standardized benchmark to evaluate AI agents on migrating enterprise Java frameworks. It tests code refactoring, d...

Jul 2, 20268 min
Local models

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

Hugging Face and Cerebras collaborate to run Gemma 4 models for real-time voice AI on local hardware, enabling low-latency speech processi...

Jul 2, 20267 min
AI research

The Next Frontier: How Artificial Intelligence is Reshaping Scientific Discovery

Artificial intelligence is revolutionizing AI research by accelerating hypothesis generation, automating experiments, and uncovering patte...

Jun 23, 20267 min
AI research

Is it agentic enough? Benchmarking open models on your own tooling

Learn how to evaluate open-source AI agents for autonomy and task completion using custom benchmarks. A practical guide for researchers an...

Jun 18, 20269 min
AI research

olmo-eval: An evaluation workbench for the model development loop

olmo-eval is an evaluation workbench designed to integrate seamlessly into the model development loop, enabling rapid iteration and system...

Jun 12, 20267 min
AI research

Can Voice Agents Handle Bilingual Customers? Benchmarking Frontier ASR on Code-Switched Speech

A clear and practical article about artificial intelligence for a professional audience.

Jun 10, 20264 min
AI agents

The Open Source Community is backing OpenEnv for Agentic RL

A clear and practical article about artificial intelligence for a professional audience.

Jun 8, 20268 min