The news is that AMD acquired Taalas, whose pitch is unusual even by AI chip standards. Instead of using general GPUs to repeatedly fetch model weights from external memory, Taalas encodes weights directly into chip structures so inference can run at very high token rates with much less memory traffic. The live demo people kept trying, chatjimmy.ai, appears to run a Llama 3.1 8B-class model and feels instant. That speed, not model quality, drove the excitement.
The useful framing that emerged is that this is not really about replacing the frontier-model race. It is about carving out a new layer of AI infrastructure for workloads where “good enough” intelligence plus extreme speed beats the newest model. People kept coming back to the same use cases: subagents inside larger systems, document and support workflows, local private inference, real-time vision or
multimodal processing, and any repetitive business task already being pushed onto cheap flash-style models. In that world, a model being six or twelve months old is not a fatal flaw if the payoff is much lower cost, much lower latency, and predictable offline behavior.
The biggest pushback was physical economics. Several commenters argued that hardwiring weights does not make the memory footprint disappear. It just changes the medium. That makes phones and other consumer devices a stretch with current designs, especially if an 8B-class model already wants a very large die and serious power. Others noted that Taalas only showed a small, older model, which underlines the gap between an impressive demo and anything close to trillion-parameter frontier systems. So the consensus landed in a narrower place than the hype. This looks plausible for specialized cards, appliances, robotics, embedded systems, and datacenter inference tiers built around mature models. It does not look like a near-term route to shipping the latest giant model in your pocket.
The mood was impressed but not gullible. People clearly saw the speed as real and potentially market-changing. They were much less convinced by fantasies about yearly phone upgrades or direct competition with state-of-the-art cloud models. The sharpest comments treated Taalas as a bet that inference is splitting into two markets: frontier systems chasing maximum capability, and frozen or slower-moving models optimized like firmware for cost, latency, and volume.