Mercury 2.5 on HotBot: What It’s Best At, Example Prompts and Limits

Hand-drawn editorial illustration clean lines warm colors A sprinter mid-stride

Most AI speed problems don’t call for a bigger model. They call for a faster one. Mercury 2.5, Inception’s newest diffusion language model, is built on that idea: instead of generating text one token at a time, it produces and refines many tokens at once. OpenRouter clocks it at roughly 1,107 tokens per second on standard GPUs (buildfastwithai.com). For workloads built on many quick model calls rather than one hard question, that number is the whole story.

This review covers what Mercury 2.5 does well, where it falls short, and how to slot it into HotBot alongside other tiers. If you want to run the prompts below as you go, you can try Mercury 2.5 on HotBot directly.

What Makes Mercury 2.5 Different

Mercury 2.5 is a diffusion large language model, or dLLM. Rather than writing tokens in sequence, it generates and refines multiple tokens in parallel, which works more like iterative editing than left-to-right composition (datacamp.com). That design is why the model chases throughput instead of frontier-level reasoning.

Inception released Mercury 2.5 as its most capable dLLM, reporting a 10-point intelligence gain over Mercury 2 and placing it alongside cost-optimized frontier models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 (businesswire.com). The update pushed the context window to 260K tokens, up from 128K, and added tunable reasoning, native tool use, JSON mode, and parallel tool calls.

One caveat matters. Mercury 2.5 is a Preview release, so its APIs, pricing, and performance can still shift (buildfastwithai.com). Treat it as a fast, cost-efficient workhorse, not a settled frontier system.

What Mercury 2.5 Is Best At

The model earns its keep when latency compounds or when an application fires off many calls. Inception and independent reviewers point to query rewriting, extraction, fast coding iteration, search ranking, support routing, and subagent tasks as its natural territory (buildfastwithai.com).

The production numbers hold up. Augment Code moved context compaction to Mercury and cut latency by 82%, from about 150 seconds to 27 seconds, while dropping cost 90% and holding quality steady; tool-search summaries came back in under a second (inceptionlabs.ai). That’s the profile to watch for: repeated, structured calls where speed and price outweigh a single hard reasoning pass.

Where it lands on quality

Mercury 2.5 sits between tiny efficient models and premium frontier systems (buildfastwithai.com). On LM Market Cap it posts an LMSYS Arena Elo of 1347 and falls under the Coding category, though the site notes it has no task benchmark data yet (lmmarketcap.com). Independent benchmark coverage remains thin, so read any quality claims as directional.

When to route elsewhere

For a single exceptionally hard reasoning task, a large frontier system is the better call (buildfastwithai.com). Model routing usually wins: let Mercury take the fast, high-volume path and pass the genuinely difficult requests to a heavier model. In HotBot, you can switch models per conversation from the full model catalog.

Mercury 2.5 Specifications at a Glance

The confirmed specs, useful when you compare Mercury 2.5 against other options:

Spec Mercury 2.5 Source
Model type Diffusion LLM (dLLM), Preview businesswire.com
Context window 260K tokens (up from 128K) businesswire.com
Throughput ~1,107 tokens/sec on standard GPUs buildfastwithai.com
Reasoning Tunable reasoning reasoncore.dev
Tool use Native tool use, parallel tool calls businesswire.com
Structured output JSON mode / schema-aligned JSON reasoncore.dev

Note what’s missing. The sources don’t confirm native image or vision capabilities for Mercury 2.5, so this guide doesn’t claim any. If your task needs vision, pick a HotBot tier that supports it.

Copy-Paste Prompts for Mercury 2.5

These prompts lean on the model’s strengths: extraction, rewriting, fast coding, and structured research. Inception’s own prompt guidance recommends leading with persona and goal, then knowledge and tools, then the current task, and putting non-negotiable constraints last, since Mercury weights recent context heavily (docs.inceptionlabs.ai).

Extraction and structured output

“Extract every company name, funding amount, and round stage from the text below. Return valid JSON as an array of objects with keys company, amount_usd, and round. If a field is missing, use null. Text: [paste text]”

“Read the support ticket below and classify it. Return JSON with category (billing, bug, feature_request, or other), priority (low, medium, high), and one_line_summary. Ticket: [paste ticket]”

Rewriting and routing

“Rewrite this search query into three alternative queries that surface different angles. Keep each under 12 words. Return them as a numbered list. Query: [paste query]”

“You are a support router. Read the message and decide which team should handle it: Accounts, Technical, or Sales. Respond with only the team name and a one-sentence reason. Message: [paste message]”

Fast coding iteration

“Here is a Python function that fails on empty input. Fix it, keep the same signature, and add two inline comments explaining the change. Return only the corrected function. Code: [paste code]”

“Convert this JSON payload into a TypeScript interface. Use precise types, mark optional fields with ?, and add no extra commentary. JSON: [paste JSON]”

Research synthesis

“Summarize the three documents below into a single brief. Structure it as: Key Findings (3 bullets), Points of Disagreement (2 bullets), and Open Questions (2 bullets). Cite which document each point comes from. Documents: [paste documents]”

For long documents that push toward the context ceiling, the 260K window beats models with smaller ones (benchlm.ai). Put the long reference material first and your active instruction at the end.

Mercury 2.5 vs HotBot’s In-House Tiers

HotBot is an independent AI chat service. One subscription covers 800+ models from every major provider, plus its own tiers: HotBot Chat, HotBot Chat Plus, HotBot Chat Pro (1M-token context, vision), and HotBot Image. Third-party model names like Mercury 2.5 just identify what’s available inside HotBot; they don’t imply any affiliation or endorsement.

How to think about the choice:

Need Best fit Why
High-volume, low-latency calls Mercury 2.5 Diffusion speed, low cost, 260K context
Very long documents (beyond 260K) HotBot Chat Pro 1M-token context window
Image understanding / vision tasks HotBot Chat Pro Vision support confirmed
Image generation HotBot Image Dedicated image engine
One hard reasoning task A frontier model in the catalog Deeper single-shot reasoning

The workflow that pays off is model routing inside a single interface. Push high-volume extraction and rewriting through Mercury 2.5, then switch to HotBot Chat Pro when a task needs a 1M-token context or vision. Everything lives in the HotBot model catalog.

Working With Your Own App Data

To have Mercury 2.5 or any HotBot model work directly with data from your apps, use HotBot connectors. Connectors are a paid-plan feature: subscribers link supported apps inside HotBot so the assistant can work with that app’s data directly. The free tier doesn’t include them.

Setup is straightforward: connect the app from inside HotBot on a paid plan. For the current list of supported apps and setup steps, check the HotBot pricing page. Pricing runs a free tier, $7.95/week, or $39.95/quarter.

The Bottom Line on Mercury 2.5

Mercury 2.5 is a fast, low-cost diffusion model built for workloads where speed and repetition matter more than a single hard reasoning pass. Its confirmed strengths cover extraction, query rewriting, fast coding iteration, search ranking, and support routing, backed by a 260K-token context window, native tool use, tunable reasoning, and JSON mode. Remember it’s a Preview model with a still-thin independent benchmark record.

The best approach isn’t Mercury 2.5 or a heavier model. It’s both. Route high-volume tasks to Mercury for speed and cost, switch to HotBot Chat Pro for 1M-token context or vision, and reach for a frontier model when a task is genuinely hard. Start testing with the prompts above by opening Mercury 2.5 on HotBot.

Frequently Asked Questions

What is Mercury 2.5 best for?

High-volume, low-latency tasks: query rewriting, data extraction, fast coding iteration, search ranking, and support routing (buildfastwithai.com). It works best where many model calls happen rather than one hard reasoning task.

What is Mercury 2.5’s context window?

Mercury 2.5 has a 260K-token context window, up from 128K in Mercury 2 (businesswire.com). For documents beyond that, HotBot Chat Pro offers a 1M-token context window.

Does Mercury 2.5 support tool use?

Yes. Mercury 2.5 adds native tool use, parallel tool calls, tunable reasoning, and JSON mode (businesswire.com). That makes it a fit for agentic and structured-output workflows.

How fast is Mercury 2.5?

OpenRouter reports throughput of about 1,107 tokens per second on standard GPUs (buildfastwithai.com). The speed comes from its diffusion architecture, which generates and refines tokens in parallel.

How much does it cost to use Mercury 2.5 online through HotBot?

One HotBot subscription covers the full model catalog, priced as a free tier, $7.95/week, or $39.95/quarter. See the HotBot pricing page for current details.

Can Mercury 2.5 handle images?

The verified sources don’t confirm native image or vision capabilities for Mercury 2.5, so this guide doesn’t claim them. For vision tasks, use HotBot Chat Pro, which supports vision.

More From hotbot.com

Gemini 3.8 Flash on HotBot: What It’s Best At, Example Prompts and Limits
HotBot Guides
Gemini 3.8 Flash on HotBot: What It’s Best At, Example Prompts and Limits
GPT-6 Astra Pro on HotBot: What It’s Best At, Example Prompts and Limits
HotBot Guides
GPT-6 Astra Pro on HotBot: What It’s Best At, Example Prompts and Limits
GPT-6 Astra on HotBot: What It’s Best At, Example Prompts and Limits
HotBot Guides
GPT-6 Astra on HotBot: What It’s Best At, Example Prompts and Limits
HotBot Image on HotBot: What It’s Best At, Example Prompts and Limits
HotBot Guides
HotBot Image on HotBot: What It’s Best At, Example Prompts and Limits
HotBot Chat Pro on HotBot: What It’s Best At, Example Prompts and Limits
HotBot Guides
HotBot Chat Pro on HotBot: What It’s Best At, Example Prompts and Limits