Test 001: The double payload
MCP outputDoes your harness pass tool output once, or twice?
Why it matters
We don’t want to flood the model with duplicate tokens. MCP tools can return the same data as text and structured output; the harness should pass it to the model once.
This test puts a different random marker in each field and asks the agent what it sees. Reporting both markers reveals duplication. Two extra probes check that text-only and structured-only responses still get through.
What counts as a pass?
Both fields sent
PASS: uses structured output
WARN: falls back to text
FAIL: duplicate, missing, or unexpected output
Text only sent
PASS: only text gets through
FAIL: missing or unexpected output
Structured only sent
PASS: only structured output gets through
FAIL: missing or unexpected output
Our rubric: prefer one structured payload, avoid duplication, preserve single-field results. This preference is not a claim about what the MCP standard mandates.
Latest approved run per harness, MCP client and version
Scroll the table sideways to compare all three responses →
| Harness | Response sent by the probe | Evidence | ||
|---|---|---|---|---|
| Both fieldscontent + structuredContent | Text onlycontent | Structured onlystructuredContent | ||
| GrokbotMCP client: CursorVersion 1.0.0Run: Oct 3, 2026 | WARNFalls back to text | PASS | FAILNo output | View ↗ |
| ChatGPT Work Cloud via WebMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 3, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| ChatGPT via WebMCP client: openai-mcpVersion 1.0.0Run: Oct 3, 2026 | PASSUses structured output | FAILNo output | PASS | View ↗ |
| Amp CLIMCP client: amp-thread-actorVersion 1Run: Oct 2, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Amp via OrbMCP client: amp-thread-actorVersion 1Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Cursor CloudMCP client: CursorVersion 1.0.0Run: Oct 2, 2026 | WARNFalls back to text | PASS | FAILNo output | View ↗ |
| OpenAI Agents APIMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Grok CLIMCP client: grok-shell-mcp-output-probeVersion 1.0.46Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Codex CLIMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| piMCP client: piVersion 1.0.0Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Cursor appMCP client: cursor-vscodeVersion 1.0.0Run: Oct 2, 2026 | WARNFalls back to text | PASS | FAILNo output | View ↗ |
| ChatGPT DesktopMCP client: openai-mcpVersion 1.0.0Run: Oct 2, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Claude CodeMCP client: claude-codeVersion 2.1.287Run: Oct 2, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Claude DesktopMCP client: Anthropic/ClaudeAIVersion 1.0.0Run: Oct 2, 2026 | WARNFalls back to text | PASS | PASS | View ↗ |
| Codex DesktopMCP client: codex-mcp-clientVersion 0.159.0-alpha.12.1Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
or unexpected output
Verdicts use reported markers. Observations describe individual sessions, not inferred model inputs.
Run it yourself
- Add the public MCP connection.
- Begin a run, then call all three probes.
- Submit the exact markers you see for review.
No client authorization is required. Connection guide. The same endpoint supports standalone manual probes.
Read the test prompt
Use the MCP output probe. Call begin_run first; supply the harness and model only if explicitly known, otherwise leave them null. Save the returned run_id, run_token, and client-declared identity. Keep run_token private; omit it from the final report and evidence. Call probe_both, probe_text_only, and probe_structured_only once each with that run_id and run_token. For each call, record the exact probe_id, text_marker, and structured_marker visible in the response. Use null for every value not visible; do not infer or invent markers. Call submit_result with the run_id, run_token, and all three reports, using those exact field names. If submission times out, retry with the same run_id, run_token, and unchanged reports. If the MCP connection reconnects, keep using the original run_id and run_token. Report the receipt, publication state, and the markers you saw; say “not visible” for null values. Do not fetch the endpoint or website separately, and do not start another run to recover a missing marker.
Method, controls & limitations
The probe generates fresh, independent random markers for content[0].text and structuredContent. An agent cannot infer an unseen marker from the other field.
probe_both returns both fields. probe_text_only returns text alone. probe_structured_only returns structured output with an empty content array. The two single-field tools are controls.
Exact random marker reproduction demonstrates visibility. A missing report alone does not establish the precise input the model received. These observations describe specific sessions, not permanent behavior of a harness or model. Unknown dates, versions and configurations stay unknown.
Reporting runs automatically compare submitted markers with privately saved probe outputs. Only reviewed submissions appear in the results. Exact token overhead was not measured.
Inspect the diagnostic probe source · Public probe landing page
Your agent here.
Run the probe and save the exact report, client/version, model, date, and evidence. Results are reviewed before publication.
Your agent submits through the public MCP connection. Connection guide →
Keep credentials and unrelated conversation content out of your report.