Samaya AI, a Mountain View-based AI platform for investment professionals, published FrontierFinance on July 29, an open benchmark designed to measure how well AI agents handle real investment workflows. The benchmark includes 220 queries and 11,543 expert-crafted rubrics covering five categories: screening and discovery, company research, sector and macro analysis, earnings and events, and coverage and catalyst monitoring. The full benchmark, evaluation code, and results are publicly available.
The results position Samaya’s proprietary system ahead of every frontier model tested. Samaya’s high-effort configuration scored 56% accuracy at roughly half the inference cost of Anthropic’s Claude Fable 5. Its low-effort mode scored 50.8% at one-quarter the cost. Among frontier models, Claude Fable 5 led at 49.2%, followed by OpenAI’s GPT-5.6 Sol at 46.8% and Claude Opus 4.8 at 45%. Open-source models including GLM 5.2 and DeepSeek V4 Pro were also evaluated, according to the company’s announcement.
Why General Benchmarks Fall Short
FrontierFinance is designed to address a gap in how AI agents are evaluated for finance. Existing finance benchmarks focus primarily on data extraction: pulling numbers from filings, summarizing earnings calls, answering factual questions. FrontierFinance tests multi-step workflows where partial accuracy is functionally useless. An agent that correctly identifies 9 out of 10 data points in a financial model but misses the one that drives the investment thesis has produced an unreliable output.
“Investment use cases are uniquely hard for AI because being almost correct is still a loss,” Samaya CEO Maithra Raghu said in the announcement. “Producing an expert-level response requires accuracy on every datapoint and every step of reasoning across a long, complex workflow, not just a plausible looking output.”
Domain-Specific Evaluation as a Trend
FrontierFinance joins a growing category of domain-specific agent benchmarks that treat general-purpose model evaluations as insufficient for real deployment decisions. The pattern is consistent across industries: teams building agent systems for specialized workflows need evaluation criteria that match the precision requirements of their domain, not aggregate scores on academic question sets. For investment professionals evaluating whether to trust an AI agent with portfolio tracking or catalyst monitoring, the difference between 45% and 56% accuracy on a domain-relevant benchmark is the difference between a useful tool and an expensive experiment.
The full performance comparison is available on Samaya’s research site.