An independent field guide to agent quirksby Polytomic
A puzzled pixel computer with a question mark on its blue screen

What the harness?

Same protocol. Different behavior.

Test 001: The double payload

MCP output

Does your harness pass tool output once, or twice?

Why it matters

We don’t want to flood the model with duplicate tokens. MCP tools can return the same data as text and structured output; the harness should pass it to the model once.

This test puts a different random marker in each field and asks the agent what it sees. Reporting both markers reveals duplication. Two extra probes check that text-only and structured-only responses still get through.

What counts as a pass?

Both fields sent

PASS: uses structured output
WARN: falls back to text
FAIL: duplicate, missing, or unexpected output

Text only sent

PASS: only text gets through
FAIL: missing or unexpected output

Structured only sent

PASS: only structured output gets through
FAIL: missing or unexpected output

Our rubric: prefer one structured payload, avoid duplication, preserve single-field results. This preference is not a claim about what the MCP standard mandates.

Observed results

Run history ↗

Latest approved run per harness, MCP client and version

Scroll the table sideways to compare all three responses →

Test 001 observations. Each result applies to one recorded session. Desired outcomes are defined in the rubric above.
HarnessResponse sent by the probeEvidence
Both fieldscontent + structuredContentText onlycontentStructured onlystructuredContent
GrokbotMCP client: CursorVersion 1.0.0Run: Oct 3, 2026WARNFalls back to textPASSFAILNo outputView ↗
ChatGPT Work Cloud via WebMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 3, 2026FAILDuplicate outputPASSPASSView ↗
ChatGPT via WebMCP client: openai-mcpVersion 1.0.0Run: Oct 3, 2026PASSUses structured outputFAILNo outputPASSView ↗
Amp CLIMCP client: amp-thread-actorVersion 1Run: Oct 2, 2026PASSUses structured outputPASSPASSView ↗
Amp via OrbMCP client: amp-thread-actorVersion 1Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
Cursor CloudMCP client: CursorVersion 1.0.0Run: Oct 2, 2026WARNFalls back to textPASSFAILNo outputView ↗
OpenAI Agents APIMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
Grok CLIMCP client: grok-shell-mcp-output-probeVersion 1.0.46Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
Codex CLIMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
piMCP client: piVersion 1.0.0Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
Cursor appMCP client: cursor-vscodeVersion 1.0.0Run: Oct 2, 2026WARNFalls back to textPASSFAILNo outputView ↗
ChatGPT DesktopMCP client: openai-mcpVersion 1.0.0Run: Oct 2, 2026PASSUses structured outputPASSPASSView ↗
Claude CodeMCP client: claude-codeVersion 2.1.287Run: Oct 2, 2026PASSUses structured outputPASSPASSView ↗
Claude DesktopMCP client: Anthropic/ClaudeAIVersion 1.0.0Run: Oct 2, 2026WARNFalls back to textPASSPASSView ↗
Codex DesktopMCP client: codex-mcp-clientVersion 0.159.0-alpha.12.1Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
PASS Desired outputWARN Text fallbackFAIL Duplicate, missing,
or unexpected output

Verdicts use reported markers. Observations describe individual sessions, not inferred model inputs.

Run it yourself

  1. Add the public MCP connection.
  2. Begin a run, then call all three probes.
  3. Submit the exact markers you see for review.
https://what-the-harness.vercel.app/mcp

No client authorization is required. Connection guide. The same endpoint supports standalone manual probes.

Read the test prompt

Use the MCP output probe. Call begin_run first; supply the harness and model only if explicitly known, otherwise leave them null. Save the returned run_id, run_token, and client-declared identity. Keep run_token private; omit it from the final report and evidence. Call probe_both, probe_text_only, and probe_structured_only once each with that run_id and run_token. For each call, record the exact probe_id, text_marker, and structured_marker visible in the response. Use null for every value not visible; do not infer or invent markers. Call submit_result with the run_id, run_token, and all three reports, using those exact field names. If submission times out, retry with the same run_id, run_token, and unchanged reports. If the MCP connection reconnects, keep using the original run_id and run_token. Report the receipt, publication state, and the markers you saw; say “not visible” for null values. Do not fetch the endpoint or website separately, and do not start another run to recover a missing marker.

Method, controls & limitations

The probe generates fresh, independent random markers for content[0].text and structuredContent. An agent cannot infer an unseen marker from the other field.

probe_both returns both fields. probe_text_only returns text alone. probe_structured_only returns structured output with an empty content array. The two single-field tools are controls.

Exact random marker reproduction demonstrates visibility. A missing report alone does not establish the precise input the model received. These observations describe specific sessions, not permanent behavior of a harness or model. Unknown dates, versions and configurations stay unknown.

Reporting runs automatically compare submitted markers with privately saved probe outputs. Only reviewed submissions appear in the results. Exact token overhead was not measured.

Inspect the diagnostic probe source · Public probe landing page

Your agent here.

Run the probe and save the exact report, client/version, model, date, and evidence. Results are reviewed before publication.

Your agent submits through the public MCP connection. Connection guide →

Keep credentials and unrelated conversation content out of your report.