Pathway Secures 0M Seed at 00M Valuation to Scale BDH Post-Transformer Architecture

AI research company Pathway has secured additional capital at a $500 million valuation, bringing its total seed funding to $30 million. The company is developing a post-Transformer architecture dubbed Baby Dragon Hatchling (BDH) designed to combine continuous in-weights adaptation, long-horizon reasoning, and memory within neural representations without relying on expanding KV caches or external retrieval pipelines. The BDH Post-Transformer Architecture Standard Transformer architectures suff

2 min
Pathway Secures 0M Seed at 00M Valuation to Scale BDH Post-Transformer Architecture

AI research company Pathway has secured additional capital at a $500 million valuation, bringing its total seed funding to $30 million. The company is developing a post-Transformer architecture dubbed Baby Dragon Hatchling (BDH) designed to combine continuous in-weights adaptation, long-horizon reasoning, and memory within neural representations without relying on expanding KV caches or external retrieval pipelines.

BDH Post-Transformer Architecture: Linear Attention and Internal Dynamic Memory

The BDH Post-Transformer Architecture

Standard Transformer architectures suffer from quadratic attention complexity relative to sequence length and separate context retention from model parameter weights. In contrast, Pathway's BDH architecture connects linear attention mechanisms with sparse key-query vectors and biologically inspired dynamic neural circuits.

According to Pathway, BDH scales contextual reasoning as a function of the model's active neuron capacity rather than a static sequence window limit. By maintaining internal state continuity, the architecture aims to support:

  1. Continuous In-Context Adaptation: The model modifies internal activation dynamics and state trajectories during runtime without requiring standard gradient fine-tuning.
  2. Linear Attention Scaling: Replacing standard O(N2)O(N^2) softmax attention matrices with linear time-complexity mechanisms reduces memory bandwidth bottlenecks during inference.
  3. Internal Temporal Reasoning: Integrating memory directly into the network dynamics enables long-horizon multi-month enterprise business processes—such as continuous quarterly financial reconciliations—without ballooning KV cache memory footprints.

Compute Partnership and Enterprise Deployment

Pathway developed and trained BDH in partnership with Amazon Web Services (AWS), utilizing Amazon SageMaker HyperPod for distributed compute and cluster orchestration.

Alongside the foundation model architecture, the startup maintains the Pathway Framework—an open-source Python engine optimized for stream processing, real-time ETL, and streaming RAG pipelines. The new funding will support increasing BDH parameter scale and training broader models targeting mathematical reasoning and benchmarks such as ARC-AGI-2 and ARC-AGI-3.

Sources

Written by

More to read

  • Prompt Compression in Production: Architecture, Latency Economics, and Degradation Trade-Offs

    As context windows expand beyond one million tokens, production LLM systems face an unexpected bottleneck: memory bandwidth and prefill latency. In high-throughput serving environments, feeding tens of thousands of tokens of few-shot demonstrations, system prompts, multi-turn conversational history, and retrieved document chunks directly into frontier models incurs heavy token costs and degrades time-to-first-token (TTFT). While early mitigation focused purely on retrieval rerankers, production

    1 min
  • MIT, Stanford, and 12 Academic Labs Launch Public AI Observatory to Track Real-World LLM Usage

    A consortium of researchers from MIT, Stanford University, and 12 other academic institutions has launched the Public AI Observatory (ai-observatory.org), an independent, auditable data repository designed to measure how individuals interact with artificial intelligence assistants in real-world settings. The initiative aims to address the empirical opacity surrounding commercial LLM deployment. While frontier AI developers such as OpenAI and Anthropic periodically release aggregated user metric

    1 min
  • DDR5 Memory Prices Climb 500% in 12 Months as AI Hyperscalers Corner Global DRAM Capacity

    Spot and contract prices for standard DDR5 dynamic random-access memory (DRAM) have climbed by up to 500 percent over the past 12 months, driven by hyper-scaler procurement teams reserving global semiconductor fabrication lines for enterprise AI accelerator memory. Historical retail and channel tracking data compiled by PCPartPicker and reported by Tom's Hardware highlights severe price spikes across high-density modules. A 128GB DDR5-6400 kit that carried an all-time low of $329 now retails fo

    1 min