inference
34 posts tagged inference.
- Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill
- Free local AI models are not charity
- Prompt lookup drafting gets faster in llama.cpp, but the win is workload-specific
- Qwen architecture rumors do not make a 3090 fast by default
- Ling 3.0 Tiny makes old CPUs interesting again
- Swift Qwen’s speed claim is a local AI reminder, measure the whole loop
- MISVO steers frozen language models at inference time
- A Small Model That Steers a Bigger One's Reasoning
- CliffCompaction makes long coding runs cheaper by refusing to summarize
- Softmax attention needs an off switch
- CXMT’s reported memory ramp matters more for inference than model hype
- LACE compresses speech tokens one codec layer at a time
- The Skill Router You Already Have: Gavel Reads Routing From a Frozen LLM
- Local LLMs are getting useful because constraints are back
- ConvMem Turns Long-Context Reasoning Into a Tree Instead of a Chain
- Declarative attention lets models skip the context they do not need
- Reasoning Without Tokens: What Soft Latent Thinking Actually Changes
- Synthetic data needs a weight limit
- llama.cpp’s CPU backlog is a map of local AI’s next gains
- Qwen 3.8 27B and the local OCR cost test
- Your Local LLM Isn't Dumb, Your Defaults Are
- A Strong Model Can Scaffold a Weak One Without Any Retraining
- BDH-CQ Makes ARC Reasoning Cheaper by Thinking in Latent Space
- A 16-node DGX Spark cluster at home: what running trillion-parameter models locally actually takes
- DeepSeek-V4-Flash on a 3090 shifts the bottleneck to DDR5
- DeepSeek-V4-Flash on a Mac is an I/O story, not a parameter-count story
- Kimi K3 on 8 GB RAM is a systems lesson, not a serving plan
- llama.cpp support is becoming the real local AI distribution layer
- The Qwen3.8 rumor is really a VRAM planning signal
- Mesh LLM makes distributed inference a networking problem
- vLLM 0.25 Deletes PagedAttention, and the Transformers Backend Catches Up
- Long-context KV caches are getting selective
- Local LLMs are becoming a workflow choice, not a hobby project
- OrbitQuant makes diffusion quantization less tied to calibration sets