Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

inference

34 posts tagged inference.

  • Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill Sep 28, 2026
  • Free local AI models are not charity Sep 27, 2026
  • Prompt lookup drafting gets faster in llama.cpp, but the win is workload-specific Sep 27, 2026
  • Qwen architecture rumors do not make a 3090 fast by default Sep 27, 2026
  • Ling 3.0 Tiny makes old CPUs interesting again Sep 26, 2026
  • Swift Qwen’s speed claim is a local AI reminder, measure the whole loop Sep 26, 2026
  • MISVO steers frozen language models at inference time Sep 25, 2026
  • A Small Model That Steers a Bigger One's Reasoning Sep 23, 2026
  • CliffCompaction makes long coding runs cheaper by refusing to summarize Sep 23, 2026
  • Softmax attention needs an off switch Sep 21, 2026
  • CXMT’s reported memory ramp matters more for inference than model hype Sep 20, 2026
  • LACE compresses speech tokens one codec layer at a time Sep 16, 2026
  • The Skill Router You Already Have: Gavel Reads Routing From a Frozen LLM Sep 15, 2026
  • Local LLMs are getting useful because constraints are back Sep 13, 2026
  • ConvMem Turns Long-Context Reasoning Into a Tree Instead of a Chain Sep 10, 2026
  • Declarative attention lets models skip the context they do not need Sep 3, 2026
  • Reasoning Without Tokens: What Soft Latent Thinking Actually Changes Sep 1, 2026
  • Synthetic data needs a weight limit Aug 31, 2026
  • llama.cpp’s CPU backlog is a map of local AI’s next gains Aug 30, 2026
  • Qwen 3.8 27B and the local OCR cost test Aug 23, 2026
  • Your Local LLM Isn't Dumb, Your Defaults Are Aug 23, 2026
  • A Strong Model Can Scaffold a Weak One Without Any Retraining Aug 13, 2026
  • BDH-CQ Makes ARC Reasoning Cheaper by Thinking in Latent Space Aug 11, 2026
  • A 16-node DGX Spark cluster at home: what running trillion-parameter models locally actually takes Aug 2, 2026
  • DeepSeek-V4-Flash on a 3090 shifts the bottleneck to DDR5 Aug 2, 2026
  • DeepSeek-V4-Flash on a Mac is an I/O story, not a parameter-count story Aug 2, 2026
  • Kimi K3 on 8 GB RAM is a systems lesson, not a serving plan Aug 2, 2026
  • llama.cpp support is becoming the real local AI distribution layer Aug 2, 2026
  • The Qwen3.8 rumor is really a VRAM planning signal Jul 19, 2026
  • Mesh LLM makes distributed inference a networking problem Jul 12, 2026
  • vLLM 0.25 Deletes PagedAttention, and the Transformers Backend Catches Up Jul 12, 2026
  • Long-context KV caches are getting selective Jul 8, 2026
  • Local LLMs are becoming a workflow choice, not a hobby project Jul 4, 2026
  • OrbitQuant makes diffusion quantization less tied to calibration sets Jul 3, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public