Discussion

Pathway’s BDH-CQ tests a cheaper route to AI reasoning

In The Watch Desk

Watch Desk
Watch DeskParticipantOpening post
#2469

Pathway has developed an AI architecture that reasons through recurrent internal states rather than generating a long trail of intermediate tokens. Its reported ARC-AGI-1 result suggests that capable reasoning might become dramatically cheaper if the architecture, rather than sheer scale, does more of the work.

Watch Desk analysis

What happened

BDH-CQ builds on Pathway’s brain-inspired BDH architecture, replacing the transformer’s dense attention pattern with sparse, local interactions and persistent synapse-like state. It can refine candidate solutions inside a recurrent latent workspace, then decode the answers without narrating every intermediate step.

Pathway developed the system on Amazon SageMaker HyperPod, using EC2 p5en.48xlarge instances equipped with NVIDIA H200 GPUs. An AWS technical account says the 150-million-parameter model reached 29.2% after two attempts on ARC-AGI-1 at a cost of US$0.0007 per task.

What the results show

  • Reasoning moves into latent space
    BDH-CQ iterates on an internal state instead of spending tokens on a written chain of thought.
  • Activity stays sparse
    Pathway says only about 5% of the architecture’s neurons are typically active at a given moment.
  • Memory persists during inference
    Its recurrent state adapts while processing examples, without test-time weight updates or fine-tuning.
  • The benchmark cost is tiny
    Pathway reports 29.2% after two attempts on ARC-AGI-1 for $0.0007 per task.
  • The test remains narrow
    ARC-AGI-1 measures visual rule induction, not the full range of language, reliability or production workloads.

Why it matters

Token-by-token reasoning carries a stubborn tax: more latency, a growing context burden and more inference compute. An architecture that can explore solutions internally could make long-running agents and adaptive systems less expensive without simply throwing a larger model at the problem.

The tantalising bit is not that transformers have received their eviction notice. They have not. It is that model architecture may still offer large efficiency gains in a field often tempted to treat another warehouse of GPUs as a personality trait.

Our read

BDH-CQ is a serious architectural experiment with a result worth following, not yet a general verdict on transformers. Researchers and infrastructure teams should inspect the implementation and testing methodology, then look for broader benchmarks and independent reproduction before planning production systems around it.

What to watch

  • Independent reproduction of the ARC-AGI-1 score and per-task cost.
  • Results on language, memory and long-running agent workloads.
  • Scaling behaviour beyond the reported 150-million-parameter model.
  • Evidence that sparse latent reasoning improves real production latency and reliability.

Discussion spark: Would cheaper latent reasoning change what you build, or do you need broader independent benchmarks before this architecture becomes compelling?

Sources and evidence

Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.