Mercury 2.5(inceptionlabs.ai)
245 points by Topfi 1 day ago | 52 comments
tl;dr: Inception has released Mercury 2.5, claimed to be the largest diffusion-based LLM ever trained, delivering 1,107 tokens/sec on NVIDIA GPUs with a 260K context window and pricing at $0.20/$0.75 per million input/output tokens (80% off at launch). The model reportedly matches cost-optimized frontier models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, with production users citing sub-200ms latencies for voice agents and 82% latency reductions for coding context compaction. Inception also previewed Mercury Voice (sub-170ms TTFT) and Mercury Router for model routing.
HN Discussion:
  • Model performs well for creative writing and speed-critical tasks like reranking
  • Disappointment that the model is not open weights despite mentioning widely available GPUs
  • ~Overeager IP-protection classifier causes errors when probing the model
  • ~Model is usable but nowhere near frontier, though compelling on price/latency
  • Enthusiasm and encouragement for pursuing diffusion-based LLM direction