inference
13 posts tagged inference.
- A Strong Model Can Scaffold a Weak One Without Any Retraining
- BDH-CQ Makes ARC Reasoning Cheaper by Thinking in Latent Space
- A 16-node DGX Spark cluster at home: what running trillion-parameter models locally actually takes
- DeepSeek-V4-Flash on a 3090 shifts the bottleneck to DDR5
- DeepSeek-V4-Flash on a Mac is an I/O story, not a parameter-count story
- Kimi K3 on 8 GB RAM is a systems lesson, not a serving plan
- llama.cpp support is becoming the real local AI distribution layer
- The Qwen3.8 rumor is really a VRAM planning signal
- Mesh LLM makes distributed inference a networking problem
- vLLM 0.25 Deletes PagedAttention, and the Transformers Backend Catches Up
- Long-context KV caches are getting selective
- Local LLMs are becoming a workflow choice, not a hobby project
- OrbitQuant makes diffusion quantization less tied to calibration sets