Agnes AI Debuts Agnes 2.5 Pro Alpha on Artificial Analysis

Agnes AI Debuts Agnes 2.5 Pro Alpha on Artificial Analysis

Agnes AI is attempting to break the economic barrier preventing developers from deploying high-reasoning agentic workflows. By launching its Agnes 2.5 Pro Alpha text model on the Artificial Analysis leaderboard, the Singapore-based firm is positioning its in-house trained models as a high-intelligence, low-cost alternative to existing frontier models that often incur prohibitive operational expenses.

Agnes 2.5 Pro Alpha Intelligence Benchmarks

The Agnes 2.5 Pro Alpha model is targeting high-reasoning tasks, specifically aiming at the "agentic" market where models must execute end-to-end workflows. On the Intelligence Index v4.1, which utilizes nine distinct evaluations, the model scored 39. This performance places it ahead of Cohere's Command A+ (23) and Mistral's Devstral 2 (19) within its specific cohort. Technical performance is further highlighted by a 67% score on Terminal-Bench v2.1 for command execution and an 88% score on GPQA Diamond, a benchmark for graduate-level science. These metrics suggest the model is designed to handle complex, search-resistant reasoning and machine-driving tasks that typically require the most expensive frontier models currently available.

Cost Efficiency and Token Restraint

A primary strategic driver for Agnes AI is reducing the "meter" costs associated with large-scale AI deployment. On the GDPval-AA v2 benchmark, the model averaged approximately three and a half cents per task. This compares to $12,300 for 10,000 tasks on GPT-5.6 Sol and $37,000 on Claude Fable 5. This massive price gap is attributed to a low token cost—$0.45 per million input tokens and $0.90 per million output tokens—and increased model restraint. While some models consume significantly more resources, Agnes 2.5 Pro Alpha utilized 26 turns and 38,000 tokens during testing. The company is offering free API access to facilitate this initial rollout.

Key Takeaways

  • Agnes 2.5 Pro Alpha scored 39 on the Intelligence Index v4.1, outperforming Cohere's Command A+ and Mistral's Devstral 2.
  • The model's pricing is set at $0.45 per million input tokens and $0.90 per million output tokens.
  • On GDPval-AA v2, the model averaged roughly $0.035 per task, significantly lower than GPT-5.6 Sol and Claude Fable 5.

TechInsyte's Take

In our view, Agnes AI is not just competing on intelligence, but on the unit economics of agency. By focusing on "restraint"—reducing the token burn per task—they are addressing the primary scaling bottleneck for enterprise agentic products. If they can maintain these performance-to-cost ratios as the model moves beyond the "Alpha" stage, they could disrupt the market for developers who are currently priced out of using top-tier frontier models for high-volume, autonomous workflows.

Questions & Answers

How does Agnes 2.5 Pro Alpha compare to competitors in reasoning benchmarks?

The model scored 39 on the Intelligence Index v4.1, which is higher than Cohere's Command A+ (23) and Mistral's Devstral 2 (19). It also achieved 88% on the GPQA Diamond science benchmark.

What is the specific pricing structure for the Agnes 2.5 Pro Alpha API?

The model is priced at $0.45 per million input tokens and $0.90 per million output tokens, which the company uses to differentiate itself from more expensive frontier models.

How does the model's task efficiency impact total cost of ownership?

The model demonstrated significant restraint, using 38,000 tokens and 26 turns per task on GDPval-AA v2. This efficiency contributed to an average task cost of approximately three and a half cents.

Is the current version of the model a finished product?

No, Agnes 2.5 Pro Alpha is the earliest version to reach Artificial Analysis and is not the finished product; the company expects it to improve in the coming months.

Source: Businesswire

TechInsyte | Technology Intelligence technology intelligence workspace

About TechInsyte | Technology Intelligence

TechInsyte is a B2B technology news and intelligence platform covering major developments across AI, cloud, cybersecurity, enterprise software, semiconductors, startups, policy, and markets. We focus on the signals that matter for decision-makers.

The idea behind TechInsyte is simple. Technology moves fast, and professionals need clear information without unnecessary noise. New platforms emerge, security risks evolve, enterprise software changes, and the AI shift continues to reshape how companies operate. We help readers understand those developments in a practical and business-focused way.

Our coverage focuses on meaningful technology updates, product launches, enterprise strategy, funding activity, regulatory change, infrastructure trends, and the broader forces shaping the technology industry. The goal is to keep every article clear, relevant, and useful for professionals who need to know what happened, why it matters, and what it could mean next.

TechInsyte is built for readers who want sharper context, cleaner coverage, and a more focused view of technology without the clutter.