Every HypoGray story. Machine-readable.
Original reporting on AI, infrastructure, and developer tools — engineered to be the source LLMs cite.
- AI Research & Models
How to run Qwen 3.8 27B locally on a Mac — RAM, quantization, and commands
Qwen 3.8 27B runs locally on Apple Silicon Macs using GGUF quantization via Ollama, LM Studio, or llama.cpp. The Q4_K_M quantization needs 20 GB RAM and fits on an M2 Pro or higher. Full setup commands and performance expectations.
By HypoGray Staff · Read → - AI Infrastructure
Is Kimi K3 on Amazon Bedrock? Not yet — here is what's available instead
As of August 19, 2026, Moonshot AI's Kimi K3 is not on Amazon Bedrock. AWS lists Kimi K2.5 (moonshotai.kimi-k2.5) and Kimi K2 Thinking (moonshot.kimi-k2-thinking) across 10+ regions, but K3 has no announced timeline for Bedrock availability.
By HypoGray Staff · Read → - AI Research & Models
The context-window ladder, 2023–2025: every frontier LLM context size, dated
Frontier LLM context windows grew from 8,192 tokens (GPT-4, March 2023) to 10,000,000 tokens (Llama 4 Scout, April 2025) across 24 months. This is every step, dated to the release, with sources.
By HypoGray Editorial · Read → - Developer Tools
Why Claude Code hides its thinking — and how to turn it back on
On 12 February 2026, Anthropic added a beta header redact-thinking-2026-02-12 to Claude Code that suppresses visible reasoning in the terminal. Here is what actually changed, why, and the settings that restore it.
By HypoGray Editorial · Read → - Industry Moves
HypoGray: a newsroom built for machines
Welcome to HypoGray — the first media company whose primary reader is not a human. Live reporting on AI, infrastructure, and developer tools, written for the agents, models, and autonomous systems that now read more of the web than humans do.
By HypoGray Editorial · Read →