Agnes AI is attempting to break the economic barrier preventing developers from deploying high-reasoning agentic workflows. By launching its Agnes 2.5 Pro Alpha text model on the Artificial Analysis leaderboard, the Singapore-based firm is positioning its in-house trained models as a high-intelligence, low-cost alternative to existing frontier models that often incur prohibitive operational expenses.
Agnes 2.5 Pro Alpha Intelligence Benchmarks
The Agnes 2.5 Pro Alpha model is targeting high-reasoning tasks, specifically aiming at the "agentic" market where models must execute end-to-end workflows. On the Intelligence Index v4.1, which utilizes nine distinct evaluations, the model scored 39. This performance places it ahead of Cohere's Command A+ (23) and Mistral's Devstral 2 (19) within its specific cohort. Technical performance is further highlighted by a 67% score on Terminal-Bench v2.1 for command execution and an 88% score on GPQA Diamond, a benchmark for graduate-level science. These metrics suggest the model is designed to handle complex, search-resistant reasoning and machine-driving tasks that typically require the most expensive frontier models currently available.
Cost Efficiency and Token Restraint
A primary strategic driver for Agnes AI is reducing the "meter" costs associated with large-scale AI deployment. On the GDPval-AA v2 benchmark, the model averaged approximately three and a half cents per task. This compares to $12,300 for 10,000 tasks on GPT-5.6 Sol and $37,000 on Claude Fable 5. This massive price gap is attributed to a low token cost—$0.45 per million input tokens and $0.90 per million output tokens—and increased model restraint. While some models consume significantly more resources, Agnes 2.5 Pro Alpha utilized 26 turns and 38,000 tokens during testing. The company is offering free API access to facilitate this initial rollout.
Key Takeaways
- Agnes 2.5 Pro Alpha scored 39 on the Intelligence Index v4.1, outperforming Cohere's Command A+ and Mistral's Devstral 2.
- The model's pricing is set at $0.45 per million input tokens and $0.90 per million output tokens.
- On GDPval-AA v2, the model averaged roughly $0.035 per task, significantly lower than GPT-5.6 Sol and Claude Fable 5.
TechInsyte's Take
In our view, Agnes AI is not just competing on intelligence, but on the unit economics of agency. By focusing on "restraint"—reducing the token burn per task—they are addressing the primary scaling bottleneck for enterprise agentic products. If they can maintain these performance-to-cost ratios as the model moves beyond the "Alpha" stage, they could disrupt the market for developers who are currently priced out of using top-tier frontier models for high-volume, autonomous workflows.
Questions & Answers
How does Agnes 2.5 Pro Alpha compare to competitors in reasoning benchmarks?
The model scored 39 on the Intelligence Index v4.1, which is higher than Cohere's Command A+ (23) and Mistral's Devstral 2 (19). It also achieved 88% on the GPQA Diamond science benchmark.
What is the specific pricing structure for the Agnes 2.5 Pro Alpha API?
The model is priced at $0.45 per million input tokens and $0.90 per million output tokens, which the company uses to differentiate itself from more expensive frontier models.
How does the model's task efficiency impact total cost of ownership?
The model demonstrated significant restraint, using 38,000 tokens and 26 turns per task on GDPval-AA v2. This efficiency contributed to an average task cost of approximately three and a half cents.
Is the current version of the model a finished product?
No, Agnes 2.5 Pro Alpha is the earliest version to reach Artificial Analysis and is not the finished product; the company expects it to improve in the coming months.
Source: Businesswire