Press Network of India

Pathway’s 150M-Parameter Model Breaks the ARC-AGI-1 Cost-Efficiency Frontier as Company Raises Funding at $500M Valuation

0 6

Pathway, an AI lab building Post-Transformer architecture and models, today published benchmark results for BDH-CQ, a 150-million-parameter reasoning model. BDH-CQ scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.0007 per task.

BDH-CQ runs approximately 11 times as cheaply per task as GPT 5.6 Luna (Low), even after OpenAI’s 80% price cut on July 30. Luna scores 34.2% against BDH-CQ’s 29.5%, representing a modest accuracy gain at approximately 11 times the cost.

ARC-AGI-1 is a public reasoning benchmark that tests whether a system can infer an underlying rule from a small number of examples and apply it to a new input, a capability often associated with human-like intelligence.

The cost gap comes from a structural difference. Many Transformer-based reasoning systems externalize their work as a chain-of-thought, generating extra tokens sequentially and feeding them back into later steps. As the trace grows, so do inference cost and latency. BDH-CQ instead performs this work in a recurrent latent state, learning from examples and refining a solution without generating an intermediate text trace. It reasons natively rather than using an intermediate-language scratchpad.

“Today’s AI pays a steep token cost for reasoning, but that cost is imposed by architecture, not by any law of intelligence. Currently, every reasoning step consumes context, adds latency, and burns compute. We show that a different architecture changes the game and opens up a whole new space in terms of how much intelligence per dollar. A 150M-parameter model, built on Pathway’s BDH architecture, reasons recurrently in latent space, and sets a new state of the art in cost efficiency on ARC-AGI-1. The bottleneck was never intelligence. It was designed,” said Zuzanna Stamirowska, CEO and Co-founder of Pathway.

Scores and comparison costs are drawn from the ARC Prize Foundation’s public leaderboard, which plots each submitted system’s benchmark score against its cost per task, as of July 2026. Pathway’s cost is computed from measured hardware time. Comparison costs are those reported to the leaderboard and may reflect API pricing for generalist models.

BDH’s ARC-AGI-1 results were evaluated by Łukasz Kaiser, a co-author of the 2017 paper that introduced the Transformer architecture, who said: “I’ve followed Pathway closely and replicated their ARC-AGI-1 results myself. Pathway shows that model architecture, not just scale, can drive the next leap in AI reasoning.”

The results were also reproduced by Richard Zhong, an NYU researcher focused on model evaluation and benchmark robustness and a co-author from Bielik.

Pathway Raises Funding at $500M Valuation

Pathway has secured additional funding at a $500 million valuation, bringing the company’s total seed funding to $30 million. The additional capital will be allocated primarily toward expanding compute capacity, including new GB300 systems.

The funding includes Id4 Ventures, TQ Ventures, Red Bridge Ventures, Kadmos Capital, and WS Investment Co., the investment arm of Wilson Sonsini, alongside angel investor Jonathan Frankle, Chief AI Scientist at Databricks.

Pathway has also appointed Adam Kurzrok as Chief Product Officer. Kurzrok previously served as a Group Product Manager for Gemini at Google DeepMind, where his work focused on model scoping, evaluation, deployment, ecosystem strategy and product development. At Pathway, he will lead product direction, including how BDH-based models are packaged, evaluated and deployed at scale.

Pathway is expanding its team as it moves into its next phase of model development and commercialization and is actively hiring across the organization.

What’s Next

Pathway plans to scale the architecture and extend the approach to more challenging reasoning benchmarks, including mathematical reasoning, ARC-AGI-2 and ARC-AGI-3, while developing a latent-reasoning Large Language Model. When its efficiency and state-tracking capabilities extend to these domains, BDH could support applications that must reason reliably as information and constraints change, from cybersecurity incident response to real-time industrial operations.

Early experiments indicate that Transformer-like scaling laws apply during pretraining at scales from 1B to 600B parameters, while preserving the latent reasoning capabilities specific to BDH-CQ.

Leave A Reply

Your email address will not be published.