xAI posted Grok 4.6 as an update over Grok 4.5, with claims of stronger benchmark performance, better coding and design output, and generous access through products like Cursor. The thread largely accepted that Grok is no longer a joke competitor. Many people who actually use frontier models said Grok 4.5 was already fast, concise, and unusually good at code and security review for the price. The recurring practical claim was not that it clearly beats Fable or Opus everywhere, but that it gets close enough that cost and speed start to dominate the decision.
The biggest caveat was that nobody really trusts benchmarks anymore. Several comments argued that recent “near-Fable” launches across labs are less mysterious than they look. Labs train continuously, release mid-training checkpoints when a rival ships, share enough techniques through hiring and social spillover, and are all riding the same wave of new compute and reinforcement learning on synthetic or verifiable tasks. That framing makes the near-simultaneous jump in model quality look more like a synchronized industry cycle than a stolen secret. It also explains why people expect narrow benchmark parity long before they grant true parity on messy long-tail work.
That gap between paper performance and lived performance came up constantly. The strongest user reports said Fable still shows better judgment on architecture and ambiguous work, while some rival models inflate their reputation by being relentless rather than wise. Grok’s perceived edge was different. People liked that it is terse, fast, and less blocked on good-faith security questions. That made it attractive as a builder or reviewer even for users who still preferred another model for planning.
The other major theme was trust. A leaked default
system prompt reinforced the view that Grok’s lighter-touch style still sits on top of standard prompt-based safety rules and likely additional filters. That sparked the usual conclusion that prompt guardrails are weak, bypassable, and mostly there for policy and liability, not as a real security boundary. Separately, many readers flatly refused to use Grok because of Musk, xAI’s handling of deepfake and child sexual abuse material controversies, and a broader fear that the vendor is politically steered and reputationally radioactive. So the practical consensus was sharp even if the mood was split: Grok 4.6 looks competitive enough that teams will test it seriously, but whether they can adopt it now depends less on model quality than on how much they care about trust, safety, and supplier risk.