You May Also Enjoy
Your LLM should answer fewer questions: a prompt-only abstention gate cut wrong answers 32%
5 minute read
Published:
CoSQ’s prompt-only abstention gate cut wrong-commitment rate 32% with 87.6% coverage — the cheapest reliability upgrade most LLM products are refusing to ship.
让你的 LLM 学会闭嘴:一个纯 prompt 的弃答门控,把错误承诺率砍了 32%
1 minute read
Published:
CoSQ 的纯 prompt 弃答门控:覆盖率 87.6% 的代价下把错误承诺率砍了 32%——这是大多数 LLM 产品拒绝上线的最便宜的可靠性升级。
DeepSeek V4.1-Flash: The Billable Unit Is No Longer the Model — It’s the Cache
6 minute read
Published:
V4.1-Flash’s 0.2pp DeepSWE ‘win’ over Opus 5 is harness noise. The substance: $0.003/M cached tokens and an 890 bytes/token KV cache that make cache-hit ratio the biggest lever in agent economics.
DeepSeek V4.1-Flash:计费单位不再是模型,而是缓存
2 minute read
Published:
V4.1-Flash 领先 Opus 5 的 0.2 个百分点只是 harness 噪声。真正的干货是 0.003 美元/M 的缓存 token 价格和 890 字节/token 的 KV 缓存——缓存命中率才是 agent 成本最大的杠杆。
