Skip to content

Alignment With Official Meta-Harness

This page documents how metaharness aligns with the official Stanford IRIS Meta-Harness release while preserving the current strengths of this repository.

Current Position

The official repository is the canonical research reference with:

  • broad domain onboarding via ONBOARDING.md
  • paper reference experiments (text classification, Terminal-Bench 2)
  • domain-specific outer loops
  • Claude Code as the shipped proposer
  • an experimental Harbor pilot (experimental/harbor_meta_harness, 2026-07-11)

This repository is a production-oriented library with:

  • packaged CLI and installable Python module
  • Codex-first proposer integration
  • filesystem-first run store and reporting
  • deterministic coding-tool benchmark workflows
  • a library kernel (optimize_harness) that hosts such as Omnigent, HarnessRouter (UHP), and AgentSky-style runtimes can embed

The right strategy is to merge strengths, not replace one with the other.

What Already Matches

  • Harness-first optimization around a fixed model surface.
  • Artifact-driven outer loop with inspectable candidate history.
  • Proposer abstraction that can support multiple providers.
  • Deterministic scoring as the decision signal for keep/discard behavior.
  • Environment bootstrap snapshots before each proposal (the main mechanism behind the paper's Terminal-Bench 2 result).
  • AHE-style change manifests on candidates.

Main Gaps To Close

  • No first-class Claude Code proposer, which is what the paper and official examples actually run.
  • No Harbor / Terminal-Bench evaluator adapter, so this repo cannot reproduce the paper headline number or later AHE / HarnessCompass results.
  • Leakage and generalization-gate checks now reject leaking diffs during search (leakage-violation); inspect surfaces the matched tokens. Keyword-dispatch heuristics beyond explicit tokens are still thin.
  • Test-time-scaling baselines (Best-of-N / sequential refine under a matched budget) are not a compare mode yet. See Wang et al. 2026.
  • Write-scope classes (prompt, skill, middleware, ...) can constrain a candidate to one class via write_scope_mode: "single-class". AHE-style per-component tracks and rollback are still not first-class.
  • Provider telemetry can still be richer for research-grade analysis.

Alignment Principles

  • Keep the current CLI and package stable.
  • Adopt official ideas as additive capabilities behind clear interfaces.
  • Preserve coding-tool workflows as a first-class domain adapter.
  • Avoid direct code vendoring from paper examples into core modules.

What's Implemented

Domain Onboarding

  • metaharness onboard <target_dir> creates ONBOARDING.md and domain_spec.md
  • Provides structured entry point for new domain work

Domain Adapter API

  • Generalized coding-tool integration into a domain adapter contract
  • Coding-tool adapter is the default built-in implementation
  • Adapter hooks for validation, search evaluation, and optional test evaluation

Split Evaluation

  • Explicit search-stage versus held-out test-stage evaluation
  • Test-stage artifacts never leak to proposer context during search
  • Run metadata fields record split definitions and leakage safeguards
  • search_result.json and optional test_result.json in run artifacts
  • Optional batch candidate proposals per iteration
  • Frontier policies beyond single scalar best, including Pareto-style policies
  • Simple hill-climb mode available as the default for low-cost workflows
  • Configurable via search_mode, proposal_batch_size, and selection_policy

Telemetry and Experiment Upgrades

  • Extended proposal telemetry with token, cost, and tool-level summaries
  • Richer experiment summary outputs for multi-objective comparisons
  • Token/tool/cost fields and expanded trial/summary columns in outputs

Leakage Gate

  • leakage_gate plus optional leakage_forbidden tokens in metaharness.json
  • Task IDs from search and held-out test tasks are forbidden in changed files
  • Rejected candidates get outcome leakage-violation; inspect and reporting count them

Write-Scope Classes

  • allowed_write_paths may be {path, class} objects as well as plain path strings
  • write_scope_mode: "single-class" rejects a candidate that touches more than one class
  • Classes: prompt, tool_desc, tool_impl, middleware, skill, subagent, memory, other

Skills Scaffold

  • metaharness scaffold coding-tool ./proj --profile skills
  • Seeds .agents/skills/<name>/SKILL.md plus Claude/Gemini skill directories and @AGENTS.md

Near-Term Roadmap

Shipped items above stay shipped. The remaining alignment work is:

  1. Claude Code proposer backend, paper-faithful, next to Codex.
  2. Harbor evaluator adapter so evaluate_search / evaluate_test can call harbor run on Terminal-Bench 2 (and later SWE-bench).
  3. Richer keyword-dispatch / test-name leakage heuristics on top of the shipped token gate. Inspired by HarnessCompass.
  4. Matched-budget sampling baseline in compare, so harness-evolution gains are not confused with extra search. Inspired by Wang et al..
  5. Deeper AHE-style component tracks and rollback, beyond single-class write-scope constraints.

Risks And Mitigations

  • Risk: overfitting core API to one research example.
  • Mitigation: keep interfaces domain-agnostic and adapter-based.

  • Risk: breaking current coding-tool user workflows.

  • Mitigation: preserve existing commands and semantics as defaults.

  • Risk: complexity jump in CLI and run layout.

  • Mitigation: gate advanced modes behind explicit flags and document layouts clearly.

  • Risk: reporting harness-evolution gains that are actually extra test-time search.

  • Mitigation: keep search/test isolation, and add matched-budget sampling baselines before claiming Harbor/TB2 numbers.

  • Risk: evolved harnesses memorizing task IDs or test names.

  • Mitigation: add a generalization-gate / regex leakage audit to inspect.

Success Criteria

  • A new domain can be scoped using onboarding files before any code changes.
  • At least one adapter can run with explicit search/test split isolation.
  • Frontier mode improves reproducibility of candidate selection in repeated trials.
  • Existing coding-tool benchmarks still run unchanged in default mode.
  • A Claude Code backend can propose against the same coding-tool benchmarks as Codex.
  • A Harbor adapter can score a candidate on a published Terminal-Bench subset without leaking held-out task names into proposer context.