fix: focus product docs and complete portable verification

This commit is contained in:
lama
2026-09-06 14:05:51 +08:00
parent 0b3f4dd431
commit 2fbf3b7867
19 changed files with 727 additions and 1189 deletions
+1
View File
@@ -0,0 +1 @@
blank_issues_enabled: true
+25
View File
@@ -0,0 +1,25 @@
## Change
Describe the user-facing problem and the smallest change that addresses it.
Link a related issue if available.
## Verification
List the commands actually run and their results. Identify skipped checks and
why they were skipped; do not mark unexecuted checks as passing.
- [ ] `python scripts/run_all_tests.py`
- [ ] If bundled sources changed: `python scripts/sync_reproduction_bundle.py --check`
- [ ] If published examples changed: verify the staged publication with `python scripts/check_publication.py --ref=`
## Compatibility and evidence
- [ ] Reviewed installation/runtime compatibility and recorded platform coverage.
- [ ] Preserved original README content/media references and linked evidence where relevant.
- [ ] Recorded changed execution assumptions, scientific meaning, or comparability; no unsupported success or reproduction claims.
- [ ] Reviewed submitted logs, profiles, snapshots, and trajectories for credentials or private data.
## Limitations
Note known gaps, migration needs, and how to undo consequential changes. Mark
checklist items that do not apply and explain briefly.
+10
View File
@@ -6,6 +6,13 @@ on:
- main
pull_request:
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs:
validate:
strategy:
@@ -16,8 +23,11 @@ jobs:
- macos-latest
- windows-latest
runs-on: ${{ matrix.os }}
timeout-minutes: 15
steps:
- uses: actions/checkout@v4
with:
persist-credentials: false
- uses: actions/setup-python@v5
with:
+37 -1
View File
@@ -2,6 +2,43 @@
## Unreleased
### September engineering hardening
#### Fixed
- Complete shared runtime and guidance in all-skills installations; preserve
the self-contained main skill and reject stale shared-runtime shadowing.
- Prevent command IDs from overwriting independent verification checks.
- Reject malformed provider responses and unsupported model parameters;
transmit supported parameters instead of silently ignoring them.
- Locate virtualenv interpreters across native Windows, MSYS2 and POSIX.
Resolve direct executables against child PATH/PATHEXT and the declared cwd;
record requested and actual argv while preserving virtualenv symlinks.
- Handle temporary-directory aliases in offline-example cleanup and installed
reference checks; preserve unexpected Git pointers rather than deleting them.
#### Added
- A small offline verification example with simulated decisions and actual
failure, preparation, process restart and independent acceptance checks.
- Installed-layout, provider, verifier, interpreter-isolation and cleanup
regressions, plus engineering acceptance criteria and operational boundaries.
- Security guidance, reproduction feedback and pull-request templates, explicit
read-only CI permissions, a job timeout and superseded-run cancellation.
#### Compatibility and evidence
- Agent state schema is `1.1`: command verdicts live under
`verification.commands`, separate from `verification.source_unchanged`.
Changed harness identity requires a fresh run; retain old checkpoints as
evidence rather than manually rewriting them to resume.
- Direct mode no longer relies on implicit host interpreter/current-directory
search. Use an explicit path or the intended child PATH. Native-shell mode
is unchanged.
- Existing public repository snapshots remain historical evidence and are not
regenerated by these changes. No successful live-model acceptance, sandbox
guarantee or paper-reproduction improvement is claimed.
### P0 / P1 reliability and agent execution
- Fix Python 3.11 benchmark cleanup and Windows Unicode/process-tree test portability.
@@ -223,4 +260,3 @@ The main flow has been trialed against a small set of public deep learning resea
- environment and asset preparation stays conservative and lightweight
- multilingual human-readable output currently focuses on English and Chinese
- the repository is intentionally not a general paper summary, benchmark design, or open-ended experiment orchestration system
+2 -2
View File
@@ -14,8 +14,8 @@ python scripts/run_all_tests.py
python scripts/check_publication.py
```
For a first hands-on exercise, run `python scripts/run_harness_lab.py` and read
the [learning roadmap](docs/PROJECT_GUIDE.md). The lab simulates model decisions
For an offline execution check, run `python scripts/run_harness_lab.py` and read
the [engineering roadmap](docs/ENGINEERING_ROADMAP.md). The example simulates model decisions
but executes real commands; it is not a model-quality benchmark. Share sanitized
reproduction feedback through the issue form; never upload credentials or
unreviewed private traces. Do not fix an acceptance failure by weakening its grader.
+145 -439
View File
@@ -1,16 +1,13 @@
# RigorPilot Skills
Run research repositories from their README, with bounded execution and auditable evidence.
RigorPilot adds section-level results without rewriting the original README.
Trusted reproduction is the default; candidate exploration requires explicit authorization.
[English](README.md) · [简体中文](README.zh-CN.md)
[![Skillselion Top 100](https://skillselion.com/badge/skills/lllllllama/rigorpilot-skills/paper-context-resolver.svg?award=1)](https://skillselion.com/skills/lllllllama/rigorpilot-skills/paper-context-resolver)
Research-first Agent Skills for Deep Learning Experiments: a harness that turns README commands into bounded runs and
auditable evidence for deep-learning repositories. Trusted reproduction is the
default; exploration requires explicit authorization.
<p>
<a href="./README.md">English</a> |
<a href="./README.zh-CN.md">简体中文</a>
</p>
<p>
<a href="https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml"><img alt="CI" src="https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml/badge.svg"></a>
<a href="https://skillselion.com/skills/lllllllama/rigorpilot-skills/ai-research-reproduction"><img alt="Listed on Skillselion" src="https://skillselion.com/badge/skills/lllllllama/rigorpilot-skills/ai-research-reproduction.svg"></a>
@@ -19,493 +16,202 @@ default; exploration requires explicit authorization.
<a href="LICENSE"><img alt="MIT License" src="https://img.shields.io/badge/license-MIT-yellow?style=flat-square"></a>
<a href="https://agentskills.io"><img alt="Agent Skills standard" src="https://img.shields.io/badge/Agent%20Skills-open%20standard-1f6feb?style=flat-square"></a>
<img alt="platforms" src="https://img.shields.io/badge/Windows%20%7C%20Linux-supported-6f42c1?style=flat-square">
<img alt="tests" src="https://img.shields.io/badge/regression-61%2F61%20passed-8250df?style=flat-square">
<img alt="external benchmark" src="https://img.shields.io/badge/external%20protocols-4%2F4%20passed-238636?style=flat-square">
<img alt="local regression" src="https://img.shields.io/badge/local%20regression-61%2F61%20passed-8250df?style=flat-square">
<a href="benchmark_outputs/external_suite_latest.json"><img alt="historical external protocols" src="https://img.shields.io/badge/historical%20protocols-4%2F4%20passed-238636?style=flat-square"></a>
</p>
<p align="center">
<a href="#examples"><strong>Examples</strong></a> ·
<a href="#evidence"><strong>Real-repo Evidence</strong></a> ·
<a href="#quick-start"><strong>Quick Start</strong></a> ·
<a href="docs/PROJECT_GUIDE.md"><strong>Learn & Roadmap</strong></a> ·
<a href="#-choose-an-entry-point"><strong>Skill Index</strong></a>
<a href="#examples"><strong>Real examples</strong></a> ·
<a href="#quick-start"><strong>Install & use</strong></a> ·
<a href="#skills"><strong>Skill index</strong></a> ·
<a href="#validation"><strong>Validation</strong></a> ·
<a href="docs/ENGINEERING_ROADMAP.md"><strong>Engineering roadmap</strong></a>
</p>
<a id="examples"></a>
## 📄 At a Glance: How RigorPilot Annotates a README
RigorPilot reads the target repository's original README directly. Every word,
blank line, and line ending stays intact; RigorPilot only inserts a status card
at the end of each section. You can see what ran, what happened, and why the
agent stopped before opening the underlying evidence.
| Original README | In-place RigorPilot verdict | Auditable evidence |
|---|---|---|
| Commands, prose, badges, images, GIFs, videos, and HTML remain unchanged | Success, partial, blocked, read-only, or authorization required | `SUMMARY.md`, `COMMANDS.md`, `LOG.md`, `status.json` |
🟢 success · 🔵 not executed · ⚪ read only · 🟡 partial · 🔴 blocked · 🟣 decision required
<div align="center">
<img src="assets/annotated-readme-preview.png" width="840" alt="Annotated README showing an error excerpt, observed metrics, evidence links, risk levels, and the training authorization boundary"/>
<br/>
<sub>Historical interface illustration: missing assets, result display and authorization boundaries. Execution provenance is not independently verified; excluded from benchmarks.</sub>
</div>
| Historical illustration (not capability evidence) | What it shows |
|---|---|
| [First attempt](examples/annotated-readme-demo/first-run/ANNOTATED_README.md) | 🟡 missing-checkpoint error display · 🟡 dataset not ready · 🟣 training awaits authorization |
| [After assets are ready](examples/annotated-readme-demo/after-setup/ANNOTATED_README.md) | 🟢 success and `mIoU` / `aAcc` display; not independently verified execution results |
### Verified on a Real Public Repository: micrograd
| Actual execution | README fidelity | Inspect directly |
|---|---|---|
| 🟢 `2` tests passed in 7.62 seconds | `8` headings = `8` annotations; stripped SHA-256 exactly matches the original file | [original repository](https://github.com/karpathy/micrograd/tree/7bc720e951fe422b8f8814aa5aa1b64121d26b4c) · [full annotated README with retained repo files](benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md) · [benchmark report](benchmark_outputs/external_micrograd.json) |
<a id="evidence"></a>
## 🧪 Real Public Repository Runs and Boundary Checks
## 📄 Real repositories, inspectable results
Original commands, prose, badges, images, videos and HTML stay in the source file.
RigorPilot splits that file into sections and inserts one evidence-linked card per section.
Removing its insertion blocks restores the retained original README byte for byte.
Each card below opens a full annotated README **beside the original README in a
retained repository checkout**. Supporting repository files are kept so relative
links and media retain their original context.
<table>
<tr>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-micrograd.png" width="100%" alt="RigorPilot micrograd reproduction showing a successful pytest run"/></a><br/>
<b>micrograd · correctness</b><br/>
<sub>🟢 2 tests passed in 7.62 s · 8 headings = 8 annotations<br/>repository files retained · SHA-256 exact</sub><br/>
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md">Open RigorPilot README inside the real repo →</a>
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-micrograd.png" width="100%" alt="micrograd: recorded pytest execution and section-level evidence"/></a><br/>
<b>micrograd · correctness checks</b><br/>
<sub>🟢 2 tests passed in 7.62 s<br/>8 headings = 8 annotations · original bytes preserved</sub><br/>
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md">Open full RigorPilot README →</a>
</td>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-mingpt.png" width="100%" alt="RigorPilot minGPT reproduction showing selected target and no unrequested execution"/></a><br/>
<b>minGPT · risk boundary</b><br/>
<sub>🔵 Test selected · no implicit model download<br/>11 headings = 11 annotations · SHA-256 exact</sub><br/>
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md">Open RigorPilot README inside the real repo →</a>
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-mingpt.png" width="100%" alt="minGPT: target selection only, with no model download or execution"/></a><br/>
<b>minGPT · selection boundary</b><br/>
<sub>🔵 Test selected, not executed · no model download<br/>11 headings = 11 annotations · original bytes preserved</sub><br/>
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md">Open full RigorPilot README →</a>
</td>
</tr>
<tr>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md"><img src="assets/showcase/external-pytorch-mnist.png" width="100%" alt="RigorPilot PyTorch MNIST reproduction showing bounded training startup"/></a><br/>
<b>PyTorch MNIST · data and metric capture</b><br/>
<sub>🟡 bounded startup · loss 0.038893<br/>1 heading = 1 annotation · SHA-256 exact</sub><br/>
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md">Open RigorPilot README inside the real repo →</a>
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md"><img src="assets/showcase/external-pytorch-mnist.png" width="100%" alt="PyTorch MNIST: partial bounded training and captured loss"/></a><br/>
<b>PyTorch MNIST · bounded startup</b><br/>
<sub>🟡 Partial training · observed loss 0.038893<br/>1 heading = 1 annotation · original bytes preserved</sub><br/>
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md">Open full RigorPilot README →</a>
</td>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-nanogpt.png" width="100%" alt="RigorPilot nanoGPT Shakespeare reproduction showing a bounded CPU training run"/></a><br/>
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-nanogpt.png" width="100%" alt="nanoGPT Shakespeare: partial CPU training and captured train and validation losses"/></a><br/>
<b>nanoGPT Shakespeare · bounded training</b><br/>
<sub>🟡 train loss 4.1676 · validation loss 4.1649<br/>11 headings = 11 annotations · SHA-256 exact</sub><br/>
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md">Open RigorPilot README inside the real repo →</a>
<sub>🟡 Partial · train loss 4.1676 · validation loss 4.1649<br/>11 headings = 11 annotations · original bytes preserved</sub><br/>
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md">Open full RigorPilot README →</a>
</td>
</tr>
</table>
<p align="center">
<a href="benchmark_outputs/EXTERNAL_REPRODUCTIONS.md"><b>Browse all four reproductions in one evidence index →</b></a>
</p>
[All four cases and upstream links](benchmark_outputs/EXTERNAL_REPRODUCTIONS.md) ·
[Recorded suite](benchmark_outputs/external_suite_latest.json) ·
[Case definitions](benchmarks/external_cases.json) · [Methodology](benchmarks/README.md)
**Recorded result:** `61/61` regression scripts and `4/4` historical external protocols
passed; the external suite took `251.0 s`, used at most `98.67 MiB` per
workspace, made `0` API calls, removed every temporary workspace, and retained
about `17.9 MiB` of tracked repository showcase snapshots.
[Latest suite JSON](benchmark_outputs/external_suite_latest.json) ·
[history](benchmark_outputs/external_suite_history.jsonl) ·
[case definitions](benchmarks/external_cases.json) ·
[methodology and limits](benchmarks/README.md)
> `partial` proves bounded startup, metric capture, timeout handling, source
> integrity, and cleanup. It does not prove convergence or reproduce a paper score.
These are historical, commit-pinned deterministic runs: **4/4 case protocols**
passed in `251.0 s`, with a peak workspace of `98.67 MiB` and `0` model API calls.
The zero-API count applies only to that suite. Selection-only and partial cases
are not completed evaluations, converged training or reproduced paper scores.
<a id="quick-start"></a>
## 🚀 Install
## 🚀 Install and use
All skills:
Install all skills:
```bash
npx skills add lllllllama/rigorpilot-skills --all
```
Only the trusted reproduction skill:
Or install only the self-contained reproduction skill:
```bash
npx skills add lllllllama/rigorpilot-skills --skill ai-research-reproduction
```
After installation, open the target repository in a Skills-capable agent:
Open the target repository in a Skills-capable agent, then ask:
> Use ai-research-reproduction. Read the original README and select the smallest documented evaluation. Stop for approval before large downloads or long training; preserve the source and record execution evidence.
> Use ai-research-reproduction: run the smallest README-documented evaluation, preserve the source and write evidence to repro_outputs/. Ask before large downloads or long training.
The main skill works alone; install all skills when using companion entrypoints.
Your existing agent loads the skill; the standalone model runner is optional.
The main skill works alone; choose **all skills** for companion and leaf entrypoints.
Your existing agent loads the skill. The standalone model runner is optional.
[Client compatibility](references/client-compatibility-policy.md)
To learn recovery from a clone of this project:
## What it does—and does not do
README → documented target → reviewed setup → bounded execution → verification → evidence.
- Preserves source meaning; records assumptions, deviations, failures and blockers.
- Records process state, logs and attempt lineage; supports explicit cancellation,
recovery and retry through the persistent runtime.
- Separates trusted reproduction from explicitly authorized, candidate-only exploration.
- Checks execution criteria independently of the model's completion claim.
This is **local execution, not an OS sandbox**. Approved commands can access the
host and network; use trusted repositories. Resource admission and between-action
budget checks are not hard OS quotas or subscription-balance monitoring.
The optional model loop currently supports Anthropic Messages and reviewed command
IDs, not unrestricted source repair. **No successful live-model acceptance is
recorded yet**: three provider attempts returned HTTP 502. Other model profiles
are metadata, not proof of working transports or equivalent model performance.
[Runner and recovery contract](skills/ai-research-reproduction/references/agent-runner.md) ·
[Implementation evidence and limits](docs/P0_P1_DELIVERY.md)
<a id="skills"></a>
## 🎯 Skill index
| Task | Skill |
|---|---|
| Reproduce from README commands | [`ai-research-reproduction`](skills/ai-research-reproduction/SKILL.md) |
| Read-only repository analysis | [`analyze-project`](skills/analyze-project/SKILL.md) |
| Prepare environment, data and weights | [`env-and-assets-bootstrap`](skills/env-and-assets-bootstrap/SKILL.md) |
| Run documented inference or evaluation | [`minimal-run-and-audit`](skills/minimal-run-and-audit/SKILL.md) |
| Start or verify training conservatively | [`run-train`](skills/run-train/SKILL.md) |
| Diagnose before proposing a patch | [`safe-debug`](skills/safe-debug/SKILL.md) |
| Coordinate authorized candidate exploration | [`ai-research-explore`](skills/ai-research-explore/SKILL.md) |
| Implement a candidate change on an isolated branch | [`explore-code`](skills/explore-code/SKILL.md) |
| Execute a bounded candidate experiment | [`explore-run`](skills/explore-run/SKILL.md) |
Two helpers support orchestration: `repo-intake-and-plan` and `paper-context-resolver`.
Exploration requires a durable `current_research` anchor and a frozen comparison
contract. Candidate results never become trusted baseline results by declaration.
[Routing](references/routing-policy.md) · [Research loop](references/research-thinking-loop.md) ·
[Campaign inputs](skills/ai-research-explore/references/research-campaign-spec.md)
## 📦 Evidence bundle
| Artifact | What to inspect |
|---|---|
| `repro_outputs/ANNOTATED_README.md` | Original README with inserted section verdicts |
| `SUMMARY.md`, `COMMANDS.md`, `LOG.md`, `status.json` | Outcome, exact commands, observations and machine-readable status |
| `PATCHES.md`, `SCIENTIFIC_CHANGELOG.md`, `COMPARABILITY_REPORT.md` | Changes, scientific meaning and comparison boundaries |
| `_runtime/<run_id>/` | Process state, events, resource samples and stdout/stderr |
| `agent_state.json`, `trajectory.jsonl` | Optional model runner's checkpoints, tool calls and reported usage |
🟢 success · 🔵 not executed · ⚪ read only · 🟡 partial · 🔴 blocked · 🟣 decision required
Ordinary output under `repro_outputs/` may change the base path of relative media
links. Published showcases and the offline example also provide a source-adjacent
`RIGORPILOT_README.md`; do not remove their retained supporting files.
[Output contract](references/output-contract.md) · [Rigor principles](references/research-rigor-principles.md)
<a id="validation"></a>
## ✅ Offline validation
From a clone of this project, with Python 3.11+ and Git:
```bash
python scripts/run_harness_lab.py
```
**Offline teaching lab: scripted decisions, real processes.** No API, GPU or model
downloads. Failure → preparation → pause → process restart → independent checks;
not evidence of live-model capability. Requires Python 3.11+ and Git.
[Project assessment & learning roadmap](docs/PROJECT_GUIDE.md) ·
[Original lab README](examples/harness-lab/README.md)
<details>
<summary>Other install paths, agent commands, and runtime controls</summary>
Optional model/tool loop: [run and resume a reviewed task](skills/ai-research-reproduction/references/agent-runner.md).
Engineering tests cover recovery and independent verification; live-provider
acceptance remains blocked by gateway HTTP 502. [P0/P1 evidence and limits](docs/P0_P1_DELIVERY.md).
Claude Code commands: `/ai-research-reproduction`, `/ai-research-explore`,
`/analyze-project`, `/safe-debug`.
Each executed command receives a run ID and writes atomic state, append-only
events, resource samples, and complete stdout/stderr under
`<output-dir>/_runtime/<run_id>/`. Cancellation, restart recovery, and explicit
retry preserve process and attempt lineage. Model profiles record identity and
capabilities without credentials.
The preferred source is `lllllllama/rigorpilot-skills`; the legacy
`lllllllama/ai-paper-reproduction-skills` slug remains a compatibility fallback.
</details>
## 📄 Output Bundle
Every run writes the original README plus section-level verdicts to
`repro_outputs/ANNOTATED_README.md`. Each verdict links to `SUMMARY.md`,
`COMMANDS.md`, `LOG.md`, and `status.json`; the header records a rubric-style
coverage score. 🟢 success · 🔵 not executed · ⚪ read only · 🟡 partial ·
🔴 blocked · 🟣 decision required.
## 🎯 Choose an Entry Point
| What you want to do | RigorPilot display name | Current skill slug |
|---|---|---|
| Reproduce a deep learning repository from README commands | Rigor Reproduce | `ai-research-reproduction` |
| Analyze repository structure, entrypoints, and risks without editing | Rigor Analyze / Audit | `analyze-project` |
| Prepare environment, datasets, weights, and cache assumptions | Rigor Setup | `env-and-assets-bootstrap` |
| Run documented inference or evaluation conservatively | Rigor Run | `minimal-run-and-audit` |
| Start or verify training conservatively | Rigor Train | `run-train` |
| Debug a failure safely, diagnose before patching | Rigor Debug / Audit | `safe-debug` |
| Explore candidates on top of `current_research` | Rigor Explore | `ai-research-explore` |
| Implement candidate changes on an isolated branch | Rigor Improve | `explore-code` |
| Run small probes or short-cycle experiments | Rigor Explore / Improve | `explore-run` |
Bundled helper skills are usually called by orchestrators:
- `repo-intake-and-plan`
- `paper-context-resolver`
## 🛣️ Lane Model
### 🔒 Trusted Lane
Use this lane for reproduction, setup, read-only analysis, conservative
execution, training verification, and safe debugging.
- Primary entrypoint: `ai-research-reproduction`
- Output directories: `repro_outputs/`, `train_outputs/`, `analysis_outputs/`, `debug_outputs/`
- Core requirement: preserve scientific meaning, minimize semantic changes, and record assumptions, blockers, and evidence.
### 🧪 Explore Lane
Use this lane only when the researcher explicitly authorizes candidate-only
exploration.
- Primary entrypoint: `ai-research-explore`
- Leaf skills: `explore-code`, `explore-run`
- Output directory: `explore_outputs/`
- Key anchor: `current_research`
`current_research` should be a durable research state such as a branch, commit,
checkpoint, run record, or already-trained local model state. Explore outputs
are always candidate results. They must not claim trusted reproduction success,
complete benchmark results, or verified novelty.
## 🔬 Core Research Principles
1. Do not chase scores blindly: score gains must have explanatory value.
2. Do not claim novelty lightly: novelty needs literature, code, or experimental evidence.
3. Do not break comparability silently: if evaluation conditions change, say that results are not directly comparable.
4. Do not disguise engineering fixes as research contributions.
5. Do not leave collaborators out of control: important changes must be auditable, reversible, and explainable.
See [references/research-rigor-principles.md](references/research-rigor-principles.md)
and [references/agent-operating-principles.md](references/agent-operating-principles.md).
## 🔁 Lifecycle View
The repository follows a shallow lifecycle-oriented routing model:
```mermaid
flowchart LR
A[Understand] --> B[Reproduce]
B --> C[Set up]
C --> D[Run or train]
D --> E[Debug]
E --> F[Report]
B -. explicit only .-> G[Explore]
G --> H[Rank candidates]
H --> F
```
The lifecycle helps the agent choose the right lane and evidence target. It
does not force every repository into a fixed implementation sequence.
## 🧠 Research Thinking Loop
Agents implement well but often think in engineering steps. Once the
researcher freezes the evaluation contract and explicitly authorizes
exploration, `ai-research-explore` runs a codified **greedy research cycle**
from observation to a fair keep-or-rollback decision
([full contract](references/research-thinking-loop.md)):
```mermaid
flowchart LR
A[Observe run evidence] --> B[Ground: paper · code · prior runs · intuition]
B --> C[Falsifiable hypothesis]
C --> D[Single-variable design]
D --> E[Bounded run]
E --> F{Fair comparison vs current best}
F -- better --> G[Keep as new candidate best]
F -- worse or unfair --> H[Roll back, record why]
G --> I[Ledger entry]
H --> I
I --> A
```
- Every hypothesis carries a **labeled evidence anchor**`paper`, `code`,
`prior-run`, or `intuition`; unanchored ideas queue in the idea bank and
never execute.
- **Greedy applies to selection, not honesty**: a keep requires comparable
evidence under the frozen contract; ties favor the simpler, cheaper change.
- Underneath: hard-gated idea ranking, atomic idea decomposition, three-layer
implementation fidelity (planned / heuristic / observed), and
executor-emitted file-level evidence.
- Lineage: adapts the greedy solution-space search of
[AIDE](https://arxiv.org/abs/2502.13138) and the managed agentic tree search
of [AI-Scientist-v2](https://arxiv.org/abs/2504.08066), constrained by
RigorPilot's comparability-first gates.
## 🌱 Continuous Learning
The shipped skills are an **immutable universal rigor core**; personalization
lives in a user-owned overlay
([policy](references/continuous-learning-policy.md)):
- Failed runs — and their later fixes — are auto-recorded as one-line lessons
in `~/.rigorpilot/lessons.jsonl` (opt out with `RIGORPILOT_LESSONS=0`).
- `python shared/scripts/lessons_store.py summarize` distills them into
`~/.rigorpilot/PERSONAL_RIGOR.md`, which skills consult at run start as the
researcher's standing preferences and known pitfalls.
- Hard rules: lessons are **advisory only** — they never relax rigor gates,
never store secrets, never edit skill files. Delete the folder and the
skills return to the universal base.
## 🧾 Suggested Research Evidence
| Artifact | Purpose |
|---|---|
| `SCIENTIFIC_CHANGELOG.md` | Records what changed, why it changed, whether it affects scientific meaning, and whether it remains comparable. |
| `COMPARABILITY_REPORT.md` | Explains whether results can still be compared to the README, paper, baseline, or SOTA reference. |
| `REPRODUCIBILITY_NOTES.md` | Records commands, configs, seeds, checkpoints, datasets, environment assumptions, and known gaps. |
| `NOVELTY_CLAIM.md` | States possible novelty as a hypothesis, with supporting evidence, missing evidence, limitations, and required ablations. |
| `ABLATION_PLAN.md` | Describes which variables must be isolated to validate a candidate change. |
| `EXPERIMENT_LEDGER.md` | Records runs, metrics, commands, artifacts, changed files, and evidence status. |
`SCIENTIFIC_CHANGELOG.md`, `COMPARABILITY_REPORT.md`, and `EXPERIMENT_LEDGER.md`
are already generated by standard trusted / explore writers. The remaining names
(`REPRODUCIBILITY_NOTES.md`, `NOVELTY_CLAIM.md`, `ABLATION_PLAN.md`) are
future-compatible evidence concepts.
## 📁 Output Directories
| Directory | Contents |
|---|---|
| `repro_outputs/` | Trusted reproduction bundle, including `ANNOTATED_README.md` |
| `train_outputs/` | Trusted training bundle |
| `analysis_outputs/` | Read-only analysis, research map, change map, eval contract, idea seeds, atomic idea map, implementation fidelity, and related outputs |
| `debug_outputs/` | Safe debug diagnosis and patch plan |
| `sources/` | Free-first research lookup records, repo-local extraction, and auditable index |
| `explore_outputs/` | Changeset, idea gate, experiment plan, manifest, ledger, candidate ranking, and related outputs |
## 🧩 Campaign Inputs
`ai-research-explore` still accepts `variant_spec.json`, but
`research_campaign.json` or `research_campaign.yaml` is preferred for Rigor
Explore campaigns.
Durable core fields:
- `current_research`
- `task_family`
- `dataset`
- `benchmark`
- `evaluation_source`
- `sota_reference`
- `compute_budget`
Optional fields:
- `candidate_ideas`
- `variant_spec`
- `research_lookup`
- `idea_policy`
- `idea_generation`
- `source_constraints`
- `feasibility_policy`
See [skills/ai-research-explore/references/research-campaign-spec.md](skills/ai-research-explore/references/research-campaign-spec.md).
## 🌐 Multi-Agent, Multi-Model
RigorPilot is model-agnostic by construction:
- **Agent Skills standard** — every skill is a spec-compliant `SKILL.md`
([agentskills.io](https://agentskills.io)), the format adopted by Claude
Code, OpenAI Codex, Cursor, VS Code, Gemini CLI, and 30+ other tools.
`npx skills add lllllllama/rigorpilot-skills` works for any of them.
- **`AGENTS.md` routing** — the root [`AGENTS.md`](AGENTS.md) gives
AGENTS.md-aware agents (Codex, Cursor, Copilot, Gemini CLI, Aider, Zed, …)
the lane model, entrypoint table, and hard rules without any install step.
- **Same contract, any model** — SKILL.md instructions carry no
model-specific tool syntax; the evidence bundles (`status.json`,
`ANNOTATED_README.md`, …) are identical whichever model executes the run,
so results stay comparable across GPT-, Claude-, and Gemini-based agents.
- **Per-skill client mirrors** — `skills/*/agents/openai.yaml` and
`.claude/commands/*` keep Codex- and Claude-specific entry points in sync
with the canonical contract.
## 🛠️ Local Install
Use the Python installer only when developing locally, needing a project-scoped
install, or manually targeting client directories.
```bash
python scripts/install_skills.py --client agents --target "$HOME/.agents/skills" --force
python scripts/install_skills.py --client codex --target "$HOME/.codex/skills" --force
python scripts/install_skills.py --client claude --target "$HOME/.claude/skills" --force
```
Project-scoped examples:
```bash
python scripts/install_skills.py --client agents --target ./.agents/skills --force
python scripts/install_skills.py --client claude --target ./.claude/skills --force
```
These commands are written to work in both Windows PowerShell and Linux shells.
## 💬 Example Prompts
**Trusted reproduction**
```text
Use ai-research-reproduction on this deep learning research repo. Stay README-first, prefer documented inference or evaluation, avoid unnecessary repo changes, and write outputs to repro_outputs/.
```
**Read-only analysis**
```text
Use analyze-project on this repo. Read the code, map the model and training entrypoints, and flag suspicious patterns without editing files.
```
**Safe debug**
```text
Use safe-debug on this traceback. Diagnose the failure first, propose the smallest safe fix, and do not patch until I approve.
```
**Candidate exploration**
```text
Use ai-research-explore with research_campaign.json. Treat the task family, dataset, evaluation source, and SOTA table as frozen inputs. Rank candidate ideas and write evidence outputs to analysis_outputs/ and explore_outputs/.
```
## ✅ Local Validation
Run everything (CI entrypoint):
This offline example uses **scripted decisions and actual processes**. It exercises
failure → preparation → pause → controller restart → independent verification,
without API calls, GPU use or model downloads. Inspect the printed `REPORT.json`
path and its linked artifacts. Existing output is never overwritten; use
`--output tmp/check-2` to repeat. It is not evidence of live-model capability.
[Example source and checks](examples/harness-lab/README.md)
Run the repository regression suite:
```bash
python scripts/run_all_tests.py
```
Basic checks:
Latest local record (2026-09-06): **61/61 scripts passed in 83.5 s**.
The CI badge links to the current Windows, Linux and macOS results.
Local tests do not substitute for live-model or held-out evaluation.
```bash
python scripts/validate_repo.py
python scripts/test_skill_registry.py
python scripts/test_trigger_boundaries.py
python scripts/test_operating_principles_structure.py
python scripts/test_claude_command_wrappers.py
python scripts/test_readme_selection.py
```
## Engineering and contributions
Core output and explore regressions:
[Engineering roadmap](docs/ENGINEERING_ROADMAP.md) · [Contributing](CONTRIBUTING.md) ·
[Security and reporting](SECURITY.md) · [CI workflow](.github/workflows/validate.yml) ·
[Reproduction feedback](https://github.com/lllllllama/RigorPilot-Skills/issues/new?template=reproduction.yml) ·
[MIT license](LICENSE)
```bash
python scripts/test_output_rendering.py
python scripts/test_readme_annotation.py
python scripts/test_train_output_rendering.py
python scripts/test_analysis_output_rendering.py
python scripts/test_safe_debug_output_rendering.py
python scripts/test_research_explore_dry_run.py
python scripts/test_research_explore_campaign_flow.py
python scripts/test_research_explore_artifact_consistency.py
python scripts/test_research_explore_variant_execution.py
python scripts/test_research_explore_nontraining_execution.py
python scripts/test_atomic_idea_decomposition.py
python scripts/test_idea_seed_generation.py
python scripts/test_implementation_fidelity.py
```
Keep acceptance checks independent, retain failed evidence and review traces
before publication. Do not publish credentials or unreviewed private repository data.
[Agent guidance](AGENTS.md) · [Operating principles](references/agent-operating-principles.md) ·
[Personalization policy](references/continuous-learning-policy.md)
Install-related regressions:
<details>
<summary>Historical interface illustration—not execution evidence</summary>
```bash
python scripts/test_bootstrap_env.py
python scripts/test_install_targets.py
python scripts/test_setup_planning.py
```
<img src="assets/annotated-readme-preview.png" width="840" alt="Historical MiniSeg interface illustration, not independently verified execution evidence"/>
## 🧭 Current Repo Snapshot
[First attempt](examples/annotated-readme-demo/first-run/ANNOTATED_README.md) ·
[After setup](examples/annotated-readme-demo/after-setup/ANNOTATED_README.md).
This older MiniSeg preview illustrates error, metric and authorization displays.
Its execution provenance is not independently verified; it is excluded from benchmarks.
- `11` skills total: `9` public skills and `2` helper skills.
- `6` trusted-lane public skills and `3` explore-lane public skills.
- `4` project-scoped Claude Code wrappers under `.claude/commands/`.
- `59` root Python scripts, including `55` test scripts.
- Documentation and command examples are kept usable from both Windows PowerShell and Linux shells.
## ⚠️ Current Limits
- The persistent queue is a single-host, single-writer scheduler; resource
requests provide admission control, not OS-level CPU, RAM, or GPU isolation.
- The external suite covers four repositories, but minGPT is selection-only
and the two training cases prove bounded startup rather than convergence or
paper-result reproduction.
- `run-train` remains a bounded training monitor; long runs must be submitted
deliberately through the queue or an external scheduler.
- Trusted reproduction avoids silent semantic changes.
- Helper skills stay narrow and are not public catch-all entrypoints.
- Exploratory work must stay isolated from trusted baselines.
- `ai-research-explore` is the governed Rigor Explore compatible slug, not an open-ended autonomous research agent.
## 📚 References
- [Research rigor principles](references/research-rigor-principles.md)
- [Deep learning experiment principles](references/deep-learning-experiment-principles.md)
- [Shared operating principles](references/agent-operating-principles.md)
- [Skill registry](references/skill-registry.json)
- [Routing policy](references/routing-policy.md)
- [Trigger boundary policy](references/trigger-boundary-policy.md)
- [Client compatibility policy](references/client-compatibility-policy.md)
- [Output contract](references/output-contract.md)
- [Research pitfall checklist](references/research-pitfall-checklist.md)
## 🧱 Scope
RigorPilot Skills is a research-first skill repository for deep learning
experiments. It focuses on scientific meaning, comparability, reproducibility,
collaborator control, and auditable workflow boundaries. It helps agents move
research forward more reliably, but it does not replace researcher judgment.
</details>
+148 -405
View File
@@ -1,120 +1,89 @@
# RigorPilot Skills
研究仓库的 README 命令转化为有界执行和可审计证据的科研 Harness
默认走可信复现;只在显式授权后进入探索。不只是更高分数,而是可验证的研究进展
研究仓库的 README 出发,执行有界任务,留下可核查的证据
RigorPilot 不重写原始 README只在各章节插入执行结果与证据链接
默认走可信复现;候选探索需要明确授权。
[English](README.md) · [简体中文](README.zh-CN.md)
[![Skillselion Top 100](https://skillselion.com/badge/skills/lllllllama/rigorpilot-skills/paper-context-resolver.svg?award=1)](https://skillselion.com/skills/lllllllama/rigorpilot-skills/paper-context-resolver)
<p>
<a href="./README.md">English</a> |
<a href="./README.zh-CN.md">简体中文</a>
</p>
<p>
<a href="https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml"><img alt="CI" src="https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml/badge.svg"></a>
<a href="https://skillselion.com/skills/lllllllama/rigorpilot-skills/ai-research-reproduction"><img alt="Listed on Skillselion" src="https://skillselion.com/badge/skills/lllllllama/rigorpilot-skills/ai-research-reproduction.svg"></a>
<a href="https://skills.sh/lllllllama/rigorpilot-skills"><img alt="skills.sh installs" src="https://skills.sh/b/lllllllama/rigorpilot-skills"></a>
<a href="https://github.com/lllllllama/RigorPilot-Skills/stargazers"><img alt="GitHub stars" src="https://img.shields.io/github/stars/lllllllama/RigorPilot-Skills?style=flat-square"></a>
<a href="LICENSE"><img alt="MIT License" src="https://img.shields.io/badge/license-MIT-yellow?style=flat-square"></a>
<a href="https://agentskills.io"><img alt="Agent Skills standard" src="https://img.shields.io/badge/Agent%20Skills-open%20standard-1f6feb?style=flat-square"></a>
<img alt="platforms" src="https://img.shields.io/badge/Windows%20%7C%20Linux-supported-6f42c1?style=flat-square">
<img alt="tests" src="https://img.shields.io/badge/regression-61%2F61%20passed-8250df?style=flat-square">
<img alt="external benchmark" src="https://img.shields.io/badge/external%20protocols-4%2F4%20passed-238636?style=flat-square">
<a href="https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml"><img alt="持续集成" src="https://github.com/lllllllama/RigorPilot-Skills/actions/workflows/validate.yml/badge.svg"></a>
<a href="https://skillselion.com/skills/lllllllama/rigorpilot-skills/ai-research-reproduction"><img alt="已收录于 Skillselion" src="https://skillselion.com/badge/skills/lllllllama/rigorpilot-skills/ai-research-reproduction.svg"></a>
<a href="https://skills.sh/lllllllama/rigorpilot-skills"><img alt="skills.sh 安装量" src="https://skills.sh/b/lllllllama/rigorpilot-skills"></a>
<a href="https://github.com/lllllllama/RigorPilot-Skills/stargazers"><img alt="GitHub 收藏量" src="https://img.shields.io/github/stars/lllllllama/RigorPilot-Skills?style=flat-square"></a>
<a href="LICENSE"><img alt="MIT 许可证" src="https://img.shields.io/badge/license-MIT-yellow?style=flat-square"></a>
<a href="https://agentskills.io"><img alt="Agent Skills 开放标准" src="https://img.shields.io/badge/Agent%20Skills-open%20standard-1f6feb?style=flat-square"></a>
<img alt="支持平台" src="https://img.shields.io/badge/Windows%20%7C%20Linux-supported-6f42c1?style=flat-square">
<img alt="本地回归" src="https://img.shields.io/badge/local%20regression-61%2F61%20passed-8250df?style=flat-square">
<a href="benchmark_outputs/external_suite_latest.json"><img alt="历史外部协议验证" src="https://img.shields.io/badge/historical%20protocols-4%2F4%20passed-238636?style=flat-square"></a>
</p>
<p align="center">
<a href="#examples"><strong>效果示例</strong></a> ·
<a href="#evidence"><strong>真实仓库证据</strong></a> ·
<a href="#quick-start"><strong>快速安装</strong></a> ·
<a href="docs/PROJECT_GUIDE.zh-CN.md"><strong>学习与路线</strong></a> ·
<a href="#-该用哪个入口"><strong>技能索引</strong></a>
<a href="#examples"><strong>真实示例</strong></a> ·
<a href="#quick-start"><strong>安装使用</strong></a> ·
<a href="#skills"><strong>技能索引</strong></a> ·
<a href="#validation"><strong>验证方式</strong></a> ·
<a href="docs/ENGINEERING_ROADMAP.zh-CN.md"><strong>工程路线</strong></a>
</p>
<a id="examples"></a>
## 📄 一眼看懂RigorPilot 如何批注 README
RigorPilot 直接读取目标仓库的原始 README保留其中每一个字、空行和换行符
只在各章节末尾插入状态卡片。你无需翻日志,就能先看清“做了什么、结果如何、
为什么停下”;需要核查时,再点击卡片里的证据链接下钻。
| 原始 README | RigorPilot 就地批注 | 可核查证据 |
|---|---|---|
| 命令、正文、徽章、图片、GIF、视频和 HTML 保持原样 | 成功、部分完成、阻塞、仅阅读或等待授权 | `SUMMARY.md``COMMANDS.md``LOG.md``status.json` |
🟢 执行成功 · 🔵 未执行 · ⚪ 仅阅读 · 🟡 部分完成 · 🔴 阻塞 · 🟣 需要决策
<div align="center">
<img src="assets/annotated-readme-preview.zh-CN.png" width="840" alt="批注版 README同时展示错误摘录、观测指标、证据链接、风险级别与训练授权边界"/>
<br/>
<sub>历史界面示意:展示缺失资源、执行结果与授权边界;执行来源未独立核验,不计入 benchmark。</sub>
</div>
| 历史界面示意(非能力证据) | 你会看到 |
|---|---|
| [首次尝试](examples/annotated-readme-demo-zh/first-run/ANNOTATED_README.md) | 🟡 缺失 checkpoint 的错误展示 · 🟡 数据未就绪 · 🟣 训练等待授权 |
| [资产就绪后](examples/annotated-readme-demo-zh/after-setup/ANNOTATED_README.md) | 🟢 成功状态与 `mIoU` / `aAcc` 指标展示;不是已核验的实跑成绩 |
### 真实公开仓库验证micrograd
| 实际执行 | 原文完整性 | 直接查看 |
|---|---|---|
| 🟢 `2` 项测试通过7.62 秒) | `8` 个标题 = `8` 条批注;剥离批注后 SHA-256 与原文件完全相同 | [原始仓库](https://github.com/karpathy/micrograd/tree/7bc720e951fe422b8f8814aa5aa1b64121d26b4c) · [保留原仓库文件的完整批注 README](benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md) · [benchmark 报告](benchmark_outputs/external_micrograd.json) |
<a id="evidence"></a>
## 🧪 真实公开仓库执行与边界验证
## 📄 真实仓库,可直接查看的结果
原始命令、正文、徽章、图片、视频和 HTML 保持不变。
RigorPilot 直接切分原文件,每个章节插入一条带证据链接的批注。
剥离全部插入块后,得到与保留的原始 README 逐字节一致的文件。
下方每张卡片都指向**保留的仓库副本中、与原 README 同目录的完整批注文件**。
相关仓库文件一并保留,相对链接和媒体仍处在原来的目录环境中。
截图内的英文来自原始仓库及其证据文件;本页说明统一使用中文。
<table>
<tr>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-micrograd.png" width="100%" alt="RigorPilot micrograd 复现,展示 pytest 成功执行"/></a><br/>
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-micrograd.png" width="100%" alt="micrograd实际测试执行与逐节证据"/></a><br/>
<b>micrograd · 正确性验证</b><br/>
<sub>🟢 2 项测试在 7.62 秒内通过 · 8 个标题 = 8 条批注<br/>原仓库文件保留 · SHA-256 完全一致</sub><br/>
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md">打开真实仓库内的 RigorPilot README →</a>
<sub>🟢 2 项测试在 7.62 秒内通过<br/>8 个标题 = 8 条批注 · 原文件字节不变</sub><br/>
<a href="benchmark_outputs/showcases/micrograd/repo/RIGORPILOT_README.md">打开完整 RigorPilot README →</a>
</td>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-mingpt.png" width="100%" alt="RigorPilot minGPT 复现,展示目标选择与未授权不执行"/></a><br/>
<b>minGPT · 风险边界</b><br/>
<sub>🔵 已选定测试 · 未隐式下载模型<br/>11 个标题 = 11 条批注 · SHA-256 完全一致</sub><br/>
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md">打开真实仓库内的 RigorPilot README →</a>
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-mingpt.png" width="100%" alt="minGPT仅选择目标没有下载或执行模型"/></a><br/>
<b>minGPT · 目标选择边界</b><br/>
<sub>🔵 已选定测试,未执行 · 未下载模型<br/>11 个标题 = 11 条批注 · 原文件字节不变</sub><br/>
<a href="benchmark_outputs/showcases/mingpt/repo/RIGORPILOT_README.md">打开完整 RigorPilot README →</a>
</td>
</tr>
<tr>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md"><img src="assets/showcase/external-pytorch-mnist.png" width="100%" alt="RigorPilot PyTorch MNIST 复现,展示有界训练启动"/></a><br/>
<b>PyTorch MNIST · 数据与指标捕获</b><br/>
<sub>🟡 有界启动 · 损失 0.038893<br/>1 个标题 = 1 条批注 · SHA-256 完全一致</sub><br/>
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md">打开真实仓库内的 RigorPilot README →</a>
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md"><img src="assets/showcase/external-pytorch-mnist.png" width="100%" alt="PyTorch MNIST:有界训练的部分结果及损失记录"/></a><br/>
<b>PyTorch MNIST · 有界启动</b><br/>
<sub>🟡 部分训练 · 观测损失 0.038893<br/>1 个标题 = 1 条批注 · 原文件字节不变</sub><br/>
<a href="benchmark_outputs/showcases/pytorch-mnist/repo/mnist/RIGORPILOT_README.md">打开完整 RigorPilot README →</a>
</td>
<td align="center" width="50%">
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-nanogpt.png" width="100%" alt="RigorPilot nanoGPT Shakespeare 复现,展示有界 CPU 训练"/></a><br/>
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md"><img src="assets/showcase/external-nanogpt.png" width="100%" alt="nanoGPT Shakespeare:部分 CPU 训练及训练和验证损失记录"/></a><br/>
<b>nanoGPT Shakespeare · 有界训练</b><br/>
<sub>🟡 训练损失 4.1676 · 验证损失 4.1649<br/>11 个标题 = 11 条批注 · SHA-256 完全一致</sub><br/>
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md">打开真实仓库内的 RigorPilot README →</a>
<sub>🟡 部分完成 · 训练损失 4.1676 · 验证损失 4.1649<br/>11 个标题 = 11 条批注 · 原文件字节不变</sub><br/>
<a href="benchmark_outputs/showcases/nanogpt-shakespeare/repo/RIGORPILOT_README.md">打开完整 RigorPilot README →</a>
</td>
</tr>
</table>
<p align="center">
<a href="benchmark_outputs/EXTERNAL_REPRODUCTIONS.zh-CN.md"><b>在一个证据索引中查看全部四项复现 →</b></a>
</p>
[四项用例与原始仓库链接](benchmark_outputs/EXTERNAL_REPRODUCTIONS.zh-CN.md) ·
[已记录的测试套件](benchmark_outputs/external_suite_latest.json) ·
[用例定义](benchmarks/external_cases.json) · [评测方法](benchmarks/README.md)
<sub>卡片内的英语来自外部仓库原始 README 与实际生成的英文证据文件;页面说明统一使用中文。</sub>
**已记录结果:** `61/61` 回归脚本、`4/4` 历史外部用例协议通过;
外部测试套件用时 `251.0 s`,单工作区最高 `98.67 MiB``0` 次 API 调用;
临时工作区全部删除,另保留约 `17.9 MiB` 的四个 tracked 仓库展示快照。
[最新测试套件 JSON](benchmark_outputs/external_suite_latest.json) ·
[历史记录](benchmark_outputs/external_suite_history.jsonl) ·
[用例定义](benchmarks/external_cases.json) ·
[方法与限制](benchmarks/README.md)
> `partial` 只证明有界启动、指标捕获、超时处理、源码完整性与清理能力;
> 不证明训练收敛,也不等于论文分数复现。
这是固定提交上的历史确定性执行记录:**4/4 用例协议**通过,
用时 `251.0 s`,工作区峰值 `98.67 MiB`,模型 API 调用 `0` 次。
零 API 仅指这次确定性套件;仅选择目标和部分训练,不代表完成评测、
训练收敛或复现论文分数。
<a id="quick-start"></a>
## 🚀 安装
## 🚀 安装使用
安装全部技能:
@@ -122,351 +91,125 @@ RigorPilot 直接读取目标仓库的原始 README保留其中每一个字
npx skills add lllllllama/rigorpilot-skills --all
```
安装可复现技能:
或仅安装可独立使用的复现技能:
```bash
npx skills add lllllllama/rigorpilot-skills --skill ai-research-reproduction
```
安装后,在目标仓库中让支持 Skills 的 Agent 执行
在支持 Skills 的代理中打开目标仓库,然后输入
> 使用 ai-research-reproduction读取原始 README,选择最小的已文档化评测。下载大资源或长训练前先停下确认;保留原文并输出执行证据
> 使用 ai-research-reproduction运行 README最小的已记录评测,保留原始源码并将证据写入 repro_outputs/。大规模下载或长训练前先确认
主技能可单独安装;使用其他叶子技能时安装全部技能。已有 Agent 加载技能,
不需要使用本项目的独立模型执行器。
主技能可单独使用其他配套入口和叶子技能请选择**安装全部技能**。
由你已有的代理加载技能,不必使用本项目的独立模型执行器。
[客户端兼容说明](references/client-compatibility-policy.md)
想先从源码了解恢复机制?克隆本项目后运行:
## 能做什么,以及边界
README → 文档目标 → 审核准备步骤 → 有界执行 → 验收 → 证据。
- 保持源码的科学含义,记录假设、偏差、失败与阻塞原因。
- 保存进程状态、日志和尝试历史;持久运行时支持显式取消、恢复与重试。
- 区分可信复现和明确授权的候选探索。
- 独立检查执行条件,不以模型声称完成作为成功依据。
这是**本机执行,不是操作系统沙箱**。已授权命令可以访问宿主机和网络,
请仅运行可信仓库。资源准入和步骤间预算检查不是系统级硬配额,
也不能监控订阅余额。
可选模型执行器目前支持 Anthropic Messages 协议与预先审核的命令编号,
不支持不受限的源码修复。**尚无成功的真实模型验收记录**
三次服务调用均返回 HTTP 502。其他模型配置仅记录元数据
不代表已实现相应服务接入或证明不同模型效果等同。
[执行与恢复契约](skills/ai-research-reproduction/references/agent-runner.md) ·
[实现证据与限制](docs/P0_P1_DELIVERY.md)
<a id="skills"></a>
## 🎯 技能索引
| 任务 | 技能 |
|---|---|
| 从 README 命令出发复现 | [`ai-research-reproduction`](skills/ai-research-reproduction/SKILL.md) |
| 只读分析仓库 | [`analyze-project`](skills/analyze-project/SKILL.md) |
| 准备环境、数据与权重 | [`env-and-assets-bootstrap`](skills/env-and-assets-bootstrap/SKILL.md) |
| 执行已记录的推理或评测 | [`minimal-run-and-audit`](skills/minimal-run-and-audit/SKILL.md) |
| 保守启动或验证训练 | [`run-train`](skills/run-train/SKILL.md) |
| 先诊断,再提出修复方案 | [`safe-debug`](skills/safe-debug/SKILL.md) |
| 协调已授权的候选探索 | [`ai-research-explore`](skills/ai-research-explore/SKILL.md) |
| 在隔离分支实现候选改动 | [`explore-code`](skills/explore-code/SKILL.md) |
| 执行有界候选实验 | [`explore-run`](skills/explore-run/SKILL.md) |
另有两个编排辅助技能:`repo-intake-and-plan``paper-context-resolver`
探索需要可追踪的 `current_research` 状态和固定的比较条件;
候选结果不能仅凭声明变成可信基线。
[路由规则](references/routing-policy.md) · [科研循环](references/research-thinking-loop.md) ·
[实验任务输入](skills/ai-research-explore/references/research-campaign-spec.md)
## 📦 输出证据包
| 文件 | 可以核查什么 |
|---|---|
| `repro_outputs/ANNOTATED_README.md` | 原始 README 与逐节插入的结论 |
| `SUMMARY.md``COMMANDS.md``LOG.md``status.json` | 结果、实际命令、观察记录与机器可读状态 |
| `PATCHES.md``SCIENTIFIC_CHANGELOG.md``COMPARABILITY_REPORT.md` | 修改、科学含义与可比性边界 |
| `_runtime/<run_id>/` | 进程状态、事件、资源采样与标准输出和错误日志 |
| `agent_state.json``trajectory.jsonl` | 可选模型执行器的检查点、工具调用与已报告用量 |
🟢 成功 · 🔵 未执行 · ⚪ 仅阅读 · 🟡 部分完成 · 🔴 阻塞 · 🟣 需要决策
常规输出位于 `repro_outputs/`,可能改变原文相对媒体链接的起始目录。
公开展示和离线示例另外提供与原文同目录的 `RIGORPILOT_README.md`
请勿删除其保留的相关仓库文件。
[输出契约](references/output-contract.md) · [科研严谨性原则](references/research-rigor-principles.md)
<a id="validation"></a>
## ✅ 离线验证
克隆本项目,安装 Python 3.11+ 和 Git 后运行:
```bash
python scripts/run_harness_lab.py
```
**离线教学实验:模拟决策真实进程** 无 API / GPU / 模型下载;演示失败→准备
暂停→跨进程恢复→独立验收,不证明真实模型能力。需 Python 3.11+ 和 Git
[学习、项目分析与改进路线](docs/PROJECT_GUIDE.zh-CN.md) ·
[实验原始 README](examples/harness-lab/README.md)
这个离线示例采用**预设决策真实进程**,验证失败 → 准备 → 暂停
控制进程重启 → 独立验收,不调用 API、不使用 GPU、不下载模型
打开命令输出中的 `REPORT.json` 及其关联证据即可核查。
已有输出不会被覆盖;再次运行可加 `--output tmp/check-2`
它不是真实模型能力证明。[示例源码与检查项](examples/harness-lab/README.md)
<details>
<summary>其他安装方式、代理命令与运行时控制</summary>
可选模型执行入口:[运行与恢复已审核任务](skills/ai-research-reproduction/references/agent-runner.md)。
工程测试已覆盖恢复与独立验证;真实模型验收因网关 HTTP 502 暂未完成。
[P0/P1 实现记录与边界](docs/P0_P1_DELIVERY.md)。
Claude Code 命令:`/ai-research-reproduction``/ai-research-explore`
`/analyze-project``/safe-debug`
每条实际命令都有 run ID并在 `<output-dir>/_runtime/<run_id>/` 下记录
原子状态、追加式事件、资源采样和完整 stdout/stderr。取消、重启恢复和
显式重试保留进程及 attempt 谱系;模型 profile 记录身份和能力,不记录凭据。
推荐源为 `lllllllama/rigorpilot-skills`
`lllllllama/ai-paper-reproduction-skills` 仅作兼容 fallback。
</details>
## 📄 输出证据包
每次运行都会把原 README 和逐节结论写入
`repro_outputs/ANNOTATED_README.md`。每条结论都可直达 `SUMMARY.md`
`COMMANDS.md``LOG.md``status.json`;文件头记录 rubric 式覆盖分。
🟢 成功 · 🔵 未执行 · ⚪ 仅阅读 · 🟡 部分完成 · 🔴 阻塞 · 🟣 需要决策。
## 🎯 该用哪个入口
| 你想做什么 | RigorPilot display name | 当前 skill slug |
|---|---|---|
| 从 README 命令出发复现深度学习仓库 | Rigor Reproduce | `ai-research-reproduction` |
| 只读分析仓库结构、入口、风险 | Rigor Analyze / Audit | `analyze-project` |
| 准备环境、数据、权重、缓存假设 | Rigor Setup | `env-and-assets-bootstrap` |
| 保守执行已记录的 inference / evaluation | Rigor Run | `minimal-run-and-audit` |
| 保守启动或验证训练 | Rigor Train | `run-train` |
| 安全调试失败,先诊断后 patch | Rigor Debug / Audit | `safe-debug` |
| 在 `current_research` 上做候选探索 | Rigor Explore | `ai-research-explore` |
| 在隔离分支实现候选改动 | Rigor Improve | `explore-code` |
| 做小样本 probe 或短周期试验 | Rigor Explore / Improve | `explore-run` |
内置 helper skills 通常由 orchestrator 调用:
- `repo-intake-and-plan`
- `paper-context-resolver`
## 🛣️ 两条主线
### 🔒 Trusted Lane
用于复现、环境准备、只读分析、保守执行、训练验证和安全调试。
- 主入口:`ai-research-reproduction`
- 输出目录:`repro_outputs/``train_outputs/``analysis_outputs/``debug_outputs/`
- 核心要求保持科学含义不变减少语义性改动显式记录假设、blocker 和证据。
### 🧪 Explore Lane
只在研究者明确授权 candidate-only 探索时启用。
- 主入口:`ai-research-explore`
- 叶子技能:`explore-code``explore-run`
- 输出目录:`explore_outputs/`
- 核心锚点:`current_research`
`current_research` 应该是可追踪的研究状态,例如 branch、commit、checkpoint、run record 或已训练模型状态。Explore 结果始终是候选结果,不能声称已经完成可信复现、完整 benchmark 或已验证 novelty。
## 🔬 科研严谨性原则
1. 不盲目追分:分数提升必须有解释价值。
2. 不轻易声称创新novelty 必须有文献、代码或实验依据。
3. 不破坏可比性:如果改变评估条件,必须说明结果不可直接比较。
4. 不隐藏工程修补:工程修补不能包装成方法贡献。
5. 不让合作者失控:重要修改必须可审计、可回滚、可解释。
详见 [references/research-rigor-principles.md](references/research-rigor-principles.md) 和 [references/agent-operating-principles.md](references/agent-operating-principles.md)。
## 🔁 生命周期视角
本仓库遵循一个浅层的、面向生命周期的路由模型:
```mermaid
flowchart LR
A[理解] --> B[复现]
B --> C[环境准备]
C --> D[运行或训练]
D --> E[调试]
E --> F[报告]
B -. 仅限显式授权 .-> G[探索]
G --> H[候选排序]
H --> F
```
生命周期帮助 agent 选择正确的 lane 和证据目标,但不会强迫每个仓库都走固定的实现顺序。
## 🧠 科研思维主循环
AI 的实现能力已经很强,但"思考"往往还停留在工程步骤上。在研究者冻结评测契约并显式授权探索后,`ai-research-explore` 执行一个成文的**贪心科研循环**——从观察出发,到一次公平的保留/回滚决策([完整契约](references/research-thinking-loop.md)
```mermaid
flowchart LR
A[观察运行证据] --> B[查证:论文 · 源码 · 历史运行 · 实验直觉]
B --> C[可证伪的假设]
C --> D[单变量实验设计]
D --> E[有界运行]
E --> F{与当前最优公平对比}
F -- 更好 --> G[保留为新的候选最优]
F -- 更差或不公平 --> H[回滚并记录原因]
G --> I[写入 ledger]
H --> I
I --> A
```
- 每个假设必须带**标注来源的证据锚点**——`paper``code``prior-run`
`intuition`;无锚点的想法进入 idea bank 排队,绝不直接执行。
- **贪心作用于选择,不作用于诚实**:保留决策必须基于冻结契约下的可比证据;
平局时偏向更简单、更便宜的改动。
- 底层机制:硬门控的 idea 排序、原子化拆解、planned / heuristic / observed
三层实现保真度、以及来自真实 executor 的文件级证据。
- 思想脉络:吸收 [AIDE](https://arxiv.org/abs/2502.13138) 的贪心解空间搜索与
[AI-Scientist-v2](https://arxiv.org/abs/2504.08066) 的受管树搜索,再用
RigorPilot 的可比性优先门槛加以约束。
## 🌱 持续学习与个性化
发布的技能是**不可变的通用科研素养内核**;个性化在用户自有的叠加层中发生([策略](references/continuous-learning-policy.md)
- 失败的运行——以及后来的修复——会自动记录为单行经验,存入
`~/.rigorpilot/lessons.jsonl``RIGORPILOT_LESSONS=0` 可关闭)。
- `python shared/scripts/lessons_store.py summarize` 将其提炼为
`~/.rigorpilot/PERSONAL_RIGOR.md`,技能在运行开始时读取,作为研究者的
长期偏好与已知坑位。
- 硬性规则:经验**仅供参考**——永不放松严谨门槛、永不存储密钥、永不修改
技能文件。删除该目录即回到通用基础版。
## 🧾 建议的科研证据体系
| Artifact | 作用 |
|---|---|
| `SCIENTIFIC_CHANGELOG.md` | 记录改了什么、为什么改、是否影响科学含义、是否仍可比较。 |
| `COMPARABILITY_REPORT.md` | 说明结果是否仍能与 README、论文、baseline 或 SOTA 参考比较。 |
| `REPRODUCIBILITY_NOTES.md` | 记录命令、配置、seed、checkpoint、数据集、环境假设和已知缺口。 |
| `NOVELTY_CLAIM.md` | 将可能创新写成假设,列出支持证据、缺失证据、限制和所需消融。 |
| `ABLATION_PLAN.md` | 说明需要隔离哪些变量才能验证候选改动。 |
| `EXPERIMENT_LEDGER.md` | 记录 run、指标、命令、artifact、变更文件和证据状态。 |
其中 `SCIENTIFIC_CHANGELOG.md``COMPARABILITY_REPORT.md``EXPERIMENT_LEDGER.md` 已由标准 trusted / explore writer 生成;其余名称(`REPRODUCIBILITY_NOTES.md``NOVELTY_CLAIM.md``ABLATION_PLAN.md`)是 future-compatible evidence concepts。
## 📁 输出目录
| 目录 | 内容 |
|---|---|
| `repro_outputs/` | trusted reproduction 输出包,含 `ANNOTATED_README.md` |
| `train_outputs/` | trusted training 输出包 |
| `analysis_outputs/` | 只读分析、research map、change map、eval contract、idea seeds、atomic idea map、implementation fidelity 等 |
| `debug_outputs/` | 安全调试诊断和 patch plan |
| `sources/` | free-first research lookup 记录、repo-local extraction 和可审计索引 |
| `explore_outputs/` | changeset、idea gate、experiment plan、manifest、ledger、candidate ranking 等 |
## 🧩 Campaign 输入
`ai-research-explore` 仍接受 `variant_spec.json`,更推荐使用 `research_campaign.json``research_campaign.yaml`
稳定核心字段:
- `current_research`
- `task_family`
- `dataset`
- `benchmark`
- `evaluation_source`
- `sota_reference`
- `compute_budget`
可选字段:
- `candidate_ideas`
- `variant_spec`
- `research_lookup`
- `idea_policy`
- `idea_generation`
- `source_constraints`
- `feasibility_policy`
详见 [skills/ai-research-explore/references/research-campaign-spec.md](skills/ai-research-explore/references/research-campaign-spec.md)。
## 🌐 多代理、多模型
RigorPilot 在设计上就是模型无关的:
- **Agent Skills 标准**——每个技能都是符合规范的 `SKILL.md`
[agentskills.io](https://agentskills.io)),该格式已被 Claude Code、
OpenAI Codex、Cursor、VS Code、Gemini CLI 等 30+ 工具采纳;
`npx skills add lllllllama/rigorpilot-skills` 对它们通用。
- **`AGENTS.md` 路由**——根级 [`AGENTS.md`](AGENTS.md) 让 AGENTS.md 感知的代理
Codex、Cursor、Copilot、Gemini CLI、Aider、Zed 等)无需安装即可获得
lane 模型、入口表和硬性规则。
- **同一契约,任意模型**——SKILL.md 指令不含任何模型特定的工具语法;
证据包(`status.json``ANNOTATED_README.md` 等)与执行模型无关,
GPT / Claude / Gemini 系代理产出的结果保持可比。
- **按客户端的镜像入口**——`skills/*/agents/openai.yaml`
`.claude/commands/*` 让 Codex 和 Claude 的专属入口与 canonical contract 同步。
## 🛠️ 本地安装
只有在本地开发、需要 project-scoped 安装,或需要手动指定客户端目录时,才建议使用 Python 安装脚本。
```bash
python scripts/install_skills.py --client agents --target "$HOME/.agents/skills" --force
python scripts/install_skills.py --client codex --target "$HOME/.codex/skills" --force
python scripts/install_skills.py --client claude --target "$HOME/.claude/skills" --force
```
项目内安装示例:
```bash
python scripts/install_skills.py --client agents --target ./.agents/skills --force
python scripts/install_skills.py --client claude --target ./.claude/skills --force
```
这些命令按 Windows PowerShell 与 Linux shell 的共同用法编写;`$HOME/...``./...` 在两类环境中都可用。
## 💬 示例提示词
**可信复现**
```text
Use ai-research-reproduction on this deep learning research repo. Stay README-first, prefer documented inference or evaluation, avoid unnecessary repo changes, and write outputs to repro_outputs/.
```
**只读分析**
```text
Use analyze-project on this repo. Read the code, map the model and training entrypoints, and flag suspicious patterns without editing files.
```
**安全调试**
```text
Use safe-debug on this traceback. Diagnose the failure first, propose the smallest safe fix, and do not patch until I approve.
```
**候选探索**
```text
Use ai-research-explore with research_campaign.json. Treat the task family, dataset, evaluation source, and SOTA table as frozen inputs. Rank candidate ideas and write evidence outputs to analysis_outputs/ and explore_outputs/.
```
## ✅ 本地自检
一键全量CI 同款入口):
运行仓库回归套件:
```bash
python scripts/run_all_tests.py
```
基础检查:
最近本地记录2026-09-06**61/61 脚本通过,用时 83.5 秒**。
持续集成徽章链接指向 Windows、Linux 和 macOS 的最新结果。
本地测试不能替代真实模型验收或未见任务评估。
```bash
python scripts/validate_repo.py
python scripts/test_skill_registry.py
python scripts/test_trigger_boundaries.py
python scripts/test_operating_principles_structure.py
python scripts/test_claude_command_wrappers.py
python scripts/test_readme_selection.py
```
## 工程与贡献
核心输出与 explore 回归:
[工程路线](docs/ENGINEERING_ROADMAP.zh-CN.md) · [贡献指南](CONTRIBUTING.md) ·
[安全与问题报告](SECURITY.md) · [持续集成配置](.github/workflows/validate.yml) ·
[提交复现反馈](https://github.com/lllllllama/RigorPilot-Skills/issues/new?template=reproduction.yml) ·
[MIT 许可证](LICENSE)
```bash
python scripts/test_output_rendering.py
python scripts/test_readme_annotation.py
python scripts/test_train_output_rendering.py
python scripts/test_analysis_output_rendering.py
python scripts/test_safe_debug_output_rendering.py
python scripts/test_research_explore_dry_run.py
python scripts/test_research_explore_campaign_flow.py
python scripts/test_research_explore_artifact_consistency.py
python scripts/test_research_explore_variant_execution.py
python scripts/test_research_explore_nontraining_execution.py
python scripts/test_atomic_idea_decomposition.py
python scripts/test_idea_seed_generation.py
python scripts/test_implementation_fidelity.py
```
保持验收独立,保留失败证据,发布前审查运行记录。
不要公开凭据或未经检查的私有仓库数据。
[代理协作说明](AGENTS.md) · [执行原则](references/agent-operating-principles.md) ·
[个性化策略](references/continuous-learning-policy.md)
安装相关回归:
<details>
<summary>历史界面示意,不作为执行证据</summary>
```bash
python scripts/test_bootstrap_env.py
python scripts/test_install_targets.py
python scripts/test_setup_planning.py
```
<img src="assets/annotated-readme-preview.zh-CN.png" width="840" alt="历史 MiniSeg 界面示意,执行来源未经独立验证"/>
## 🧭 当前仓库快照
[首次尝试](examples/annotated-readme-demo-zh/first-run/ANNOTATED_README.md) ·
[准备完成后](examples/annotated-readme-demo-zh/after-setup/ANNOTATED_README.md)。
这份旧 MiniSeg 预览只展示错误、指标和授权状态的呈现方式。
其执行来源没有独立核验,不计入基准测试。
- 共 `11` 个 skill其中 `9` 个 public skill`2` 个 helper skill。
- 共 `6` 个 trusted-lane public skill`3` 个 explore-lane public skill。
- `.claude/commands/` 下提供 `4` 个项目级 Claude Code wrappers。
- 根目录 `scripts/` 下共有 `59` 个 Python 脚本,其中 `55` 个是测试脚本。
- 文档和命令示例兼顾 Windows PowerShell 与 Linux shell。
## ⚠️ 当前限制
- 持久任务队列是单机、单写入者调度器;资源请求用于准入控制,不等于操作系统级 CPU、内存或 GPU 隔离。
- 外部 suite 覆盖四个仓库,但 minGPT 仅做选择验证,两个训练用例只证明有界启动,不证明收敛或论文结果复现。
- `run-train` 仍是受限训练监控器;长任务必须显式提交到该队列或外部调度器。
- trusted reproduction 避免静默语义改动。
- helper skills 保持窄职责,不作为公共兜底入口。
- exploratory work 必须与 trusted baseline 隔离。
- `ai-research-explore` 是受治理的 Rigor Explore 兼容 slug不是开放式 autonomous research agent。
## 📚 参考文档
- [Research rigor principles](references/research-rigor-principles.md)
- [Deep learning experiment principles](references/deep-learning-experiment-principles.md)
- [Shared operating principles](references/agent-operating-principles.md)
- [Skill registry](references/skill-registry.json)
- [Routing policy](references/routing-policy.md)
- [Trigger boundary policy](references/trigger-boundary-policy.md)
- [Client compatibility policy](references/client-compatibility-policy.md)
- [Output contract](references/output-contract.md)
- [Research pitfall checklist](references/research-pitfall-checklist.md)
## 🧱 仓库定位
RigorPilot Skills 是面向深度学习实验的科研优先 skill 仓库。它关注科学含义、可比性、可复现性、协作者可控性和可审计边界;它帮助 agent 更可靠地推进研究工作,但不替代研究者判断。
</details>
+41
View File
@@ -0,0 +1,41 @@
# Security policy
## Execution boundary
RigorPilot executes programs on the local host. It is not an operating-system
sandbox: an approved command can access files, processes, and the network with
the permissions of the user running it. Use trusted repositories and review
commands, dependency installation, and task files before execution. Use an
independently configured isolated environment for untrusted code.
The model-driven runner restricts its tools to reviewed command IDs, but this
does not restrict what those programs can do. Credential environment filtering
is a defense in depth, not a guarantee that a program cannot find credentials
elsewhere on the host. Timeouts and between-action output checks are not hard
disk, memory, CPU, or network isolation.
## Credentials and evidence
- Keep credentials in the documented environment variables, never in task or
model-profile JSON, source files, command arguments, or submitted examples.
- Review logs, prompts, model responses, repository snapshots, and trajectories
before sharing them. They may contain private code, personal paths, dataset
contents, or secrets emitted by a program; automatic complete redaction is
not provided.
- If a credential is exposed, revoke or rotate it with its provider. Removing a
file or commit does not invalidate a credential or remove existing copies.
- Do not submit private repositories or unreviewed evidence bundles in public
issues or pull requests. Prefer a minimal sanitized reproducer.
## Reporting a vulnerability
If private vulnerability reporting is enabled for this repository, use
**Report a vulnerability** on its [GitHub Security page](https://github.com/lllllllama/RigorPilot-Skills/security).
If that option is unavailable, open a minimal issue requesting a private
reporting channel, without exploit details, sensitive data, or credentials.
No private email address or response-time guarantee is currently published.
Include the affected commit, installation method, platform/Python version,
security impact, and the smallest safe reproduction. Distinguish a failure of
a documented boundary from the expected host access of explicitly approved
programs. Do not test against other people's systems or data.
+98
View File
@@ -0,0 +1,98 @@
# Engineering roadmap and acceptance criteria
[简体中文](ENGINEERING_ROADMAP.zh-CN.md) · [README](../README.md) · [Implementation record](P0_P1_DELIVERY.md)
Updated 2026-09-06. Planned work is not an implemented capability.
## Product scope
RigorPilot turns research-repository README targets into bounded execution and
auditable evidence. It targets small inference/evaluation runs, reproduction
preflight, conservative training startup and diagnosis. Trusted execution is
the default; source exploration needs explicit authorization. It does not
replace researcher judgment or change algorithms/budgets to manufacture success.
## Current capabilities
| Area | Implementation | Boundary |
|---|---|---|
| Installation | Self-contained main skill; shared bundled runtime and guides for all-skills installs | Tests cover installed layouts, 20 public CLIs and actual short execution, not a live third-party installation service |
| Execution | Processes, timeout/cancel, events/logs and explicit executable identity | Local host, not an OS sandbox; sampling/admission is not a hard resource quota |
| Recovery | Checkpoints, completed-result reuse, uncertain-dispatch blocking | No blind request replay or training-checkpoint restoration |
| Verification | Independent command checks and source integrity | The model loop's exit/stdout criteria are not paper-result matching |
| README | Byte-preserving incremental annotations with a strip round trip | Ordinary subdirectory output may break relative media context; showcases retain a source-adjacent copy |
| Models | Anthropic Messages tools, validated parameters and usage accounting | Three real attempts returned 502; no successful live acceptance. Other profile metadata does not imply transport support |
| External evidence | Four historical, commit-pinned protocols with retained source files/media | Includes selection-only and partial runs, not four paper reproductions or an unseen-task success rate |
## Acceptance layers
1. **Engineering regression:** `python scripts/run_all_tests.py` covers installed
layouts, invalid model responses, recovery, source fidelity and a real target
command asserting it executes inside the created virtualenv.
2. **Offline verification:** `python scripts/run_harness_lab.py` uses fixed
simulated decisions with real failure, preparation, pause and process restart.
No API/GPU/downloads; not model-quality evidence.
3. **Repository protocols:** [Pinned cases](../benchmarks/README.md) distinguish
selection, execution, partial completion and metric matching.
4. **Live-model acceptance:** one bounded
[micrograd canary](../benchmarks/run_agent_canary.py), only with a working
service and confirmed budget. Preserve actual model/tool traces, usage,
independent verdict and source hash before expanding the matrix.
## Delivery priorities
| Priority | Deliverable | Acceptance gate |
|---|---|---|
| P0: ongoing | Installation, portability, publication, feedback and security documentation | Installed files work; three-platform CI passes; failures are not reported as success |
| P1 | One real-model run, with no manually substituted trajectory | Responses, tools, usage and verifier evidence; stop and retain service failures |
| P1 | Optional source-adjacent annotation for ordinary runs | Nested README/media/evidence links work; original bytes and files are not overwritten |
| P2 | Frozen tasks, independent graders and same-condition baselines | Separate task completion, false success, incorrect blocking, cost, interventions and evidence integrity |
| P2 | Optional isolated executor, network/file boundaries and resource limits | Explicit threat model and boundary tests; no sandbox claim when unconfigured |
| P3 | Model regression, releases, compatibility notes and failure classification | Each version has regression evidence and change notes; historical evidence remains inspectable |
Defer large training runs, arbitrary source repair, multi-agent orchestration
and long-term memory infrastructure until demonstrated failures justify them.
## Reusable evaluation protocol (planned)
Begin with six tasks and three conditions to audit graders, then expand to
twelve frozen tasks across 46 repositories, with three repeats for key tasks.
Separate development from holdout data; existing public cases are regressions,
not unseen tasks.
A: same model with generic task instructions; B: A plus skill instructions;
C: B plus durable execution/recovery/evidence mechanics. Freeze model revision,
tool permissions, reviewed commands, budgets, source/environment and grader.
Retain raw traces for all arms. A baseline must not fail a business task merely
because it lacks a polished report. Different tools/permissions constitute an
end-to-end product comparison, not a single-factor ablation. Reviewed command
sets do not establish autonomous target discovery.
Cover normal execution, missing assets, exit-zero/wrong-result, interrupted
controller, premature completion and unauthorized large downloads. Label
injected faults separately. Protect external graders from agent modification;
check actual artifacts, metric tolerances and experimental conditions.
Record task/split/commit, harness/prompt/grader hashes, requested/returned model,
parameters, dependencies/cache, repeat, interventions, claimed/verified outcome,
trace, usage and latency. Mark unknown costs; include failed attempts in cost
per success. Report provider failures separately without removing them from
user-facing success denominators. Report both safe and incorrect blocking;
small samples warrant per-case evidence and uncertainty, not broad claims.
See [agent evaluation guidance](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents).
## Upgrades and operation
Protocol/usage tests → one live canary → paired holdout comparison → release.
Change one factor at a time and preserve old results. Use ablations to remove
obsolete scaffolding as models improve. Borrow acceptance contracts and
independent checks from [long-running harness practice](https://www.anthropic.com/engineering/harness-design-long-running-apps),
and the checkpoint/memory distinction from [persistence design](https://docs.langchain.com/oss/python/langgraph/persistence),
without copying their architectural scale.
Prefer single-case, serial, budgeted validation. Token/time gates cannot read
subscription balances or replace provider-side spending caps. Do not download
large models implicitly or replay requests with unknown outcomes. Review traces
before publication; see [security](../SECURITY.md) and [contributing](../CONTRIBUTING.md).
Remote branch protection, private vulnerability reporting and account settings
require maintainer confirmation; repository templates do not enable them.
+90
View File
@@ -0,0 +1,90 @@
# 工程路线与验收标准
[English](ENGINEERING_ROADMAP.md) · [首页](../README.zh-CN.md) · [实现记录](P0_P1_DELIVERY.md)
更新于 2026-09-06。本文区分已实现能力、验证范围和待交付工作路线不代表已完成。
## 产品范围
RigorPilot 把研究仓库 README 中的运行目标转化为有界执行和可审计证据。
适用于小型推理/评测、复现前检查、保守训练启动核验和故障分析。
可信执行是默认路径;源码探索需要明确授权。它不替代研究者判断,
也不以自由修改算法或扩大计算预算来换取成功状态。
## 当前工程能力
| 能力 | 已有实现 | 验证边界 |
|---|---|---|
| 安装 | 主技能自包含;全量安装共用 bundled 运行时和参考文件 | 测试安装后目录、20 个公开 CLI 和真实短命令;不等于实时验证第三方安装服务 |
| 执行 | 真实子进程、超时/取消、事件和日志;明确实际解释器路径 | 本机运行,不是 OS 沙箱;资源采样/准入不是硬配额 |
| 恢复 | 任务检查点、已完成命令复用、不确定派发阻塞 | 不自动重复未知结果的请求;不能恢复模型训练内部 checkpoint |
| 完成判定 | 命令验收与源码完整性分开;模型声明不能直接变为成功 | 模型循环仍主要采用退出码/输出字串检查,不等于论文指标匹配 |
| README | 原始字节分块后增量插入;剥离插入块可还原原文件 | 普通子目录输出未必保留相对媒体显示上下文;公开展示另外提供源旁副本 |
| 模型 | Anthropic Messages 工具循环、参数校验、用量记录 | 三次真实调用返回 502尚无成功 live 验收;其他 provider 的 profile 不等于调用支持 |
| 外部证据 | 固定 commit 的四项历史用例,保留原仓库文件和媒体 | 包括选择阶段和部分训练;不证明四篇论文复现或未知仓库成功率 |
## 验收层次
1. **结构与回归**`python scripts/run_all_tests.py`。包括完整/单技能/目录安装、
异常模型响应、恢复、原文完整性和真实虚拟环境内的目标命令。
2. **离线验证示例**`python scripts/run_harness_lab.py`。固定模拟决策驱动真实
失败、准备、暂停和跨进程恢复;不代表模型能力,无 API/GPU/下载。
3. **真实仓库协议**[固定版本用例](../benchmarks/README.md)。明确选择阶段、
执行成功、部分完成与指标匹配,不能把它们合并成论文复现率。
4. **真实模型验收**:只在服务可用、预算已确认时运行一个
[micrograd canary](../benchmarks/run_agent_canary.py)。保留实际模型/工具轨迹、
用量、独立判定和源码哈希后,才进入更大的矩阵。
## 交付优先级
| 优先级 | 交付项 | 完成门槛 |
|---|---|---|
| P0持续维护 | 安装、跨平台执行、证据发布、问题反馈与安全说明 | 安装后的文件可用;三平台 CI 通过;错误不被包装成成功 |
| P1真实调用闭环 | 一例真实模型执行;无人工替写轨迹 | 模型回复、工具调用、用量和验收结果齐全;服务失败如实保留并停止 |
| P1普通输出可浏览 | 将源旁批注副本推广为普通入口的可选输出 | 嵌套 README、相对图像和证据链接有效不改原文、不覆盖原文件 |
| P2效果评测 | 冻结任务、独立评分器、同条件基线 | 区分业务完成、误报、错误阻塞、成本、干预和证据完整性 |
| P2执行隔离 | 可选隔离执行后端、网络/文件边界及资源限制 | 威胁模型明确;越界和超限测试通过;未配置时不声称有沙箱 |
| P3维护与升级 | 模型回归、版本发布、兼容说明、失败分类 | 新版本有回归结果与变更说明,历史证据仍可核查 |
不优先扩展大规模训练、任意源码修补、多代理编排或长期记忆设施。
新增机制应针对已复现的问题,并能通过验收证明必要性。
## 长期评测协议(待实现)
先用 6 个任务、3 个条件的小规模试验审查评分器,再扩展为 12 个冻结任务、
46 个仓库,关键任务重复 3 次。开发集与保留集分开;现有公开案例属于
开发/回归材料,不能冒充未知任务。
A同模型与通用任务提示BA 加技能指令CB 加持久化执行、恢复和证据机制。
固定模型版本、工具权限、审核命令、预算、源代码、环境和评分器。所有组保留
原始轨迹,不能因基线没有漂亮报告就判业务失败。权限或工具不同的比较必须
标为端到端产品对照,不解释为单一机制的收益。预审命令也不证明自主发现目标能力。
任务覆盖正常执行、缺资源准备重试、退出码成功但结果错误、控制器中断、
提前宣告成功、需要未授权的大额下载。真实任务与注入故障分开标记。
评分器应检查实际产物、指标容差和实验条件,并与执行中的 Agent 隔离。
记录 task/split、repo commit、harness/prompt/grader hash、模型请求名/返回名、
参数、依赖与缓存条件、重复编号、人工干预、自报结果、独立判定、轨迹、用量和耗时。
费用未知须标未知包含失败尝试的成本。provider 故障单列,但不从用户成功率
的分母中删除。同时报告安全阻塞和错误阻塞,防止靠全部拒绝取得虚假的好结果。
小样本先报告逐例结果和不确定性,不作泛化承诺。
参考 [Agent 评测方法](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)。
## 模型与运行时升级
协议/用量契约测试 → 一个真实 canary → 同条件保留集对照 → 发布。
一次只改变模型、提示词或工具中的一个因素;保留旧结果。模型进步后,
通过消融实验决定是否删除旧补救规则,而不是永久积累约束。
借鉴 [长任务 Harness 实践](https://www.anthropic.com/engineering/harness-design-long-running-apps)
中的验收契约与独立检查,以及 [持久化设计](https://docs.langchain.com/oss/python/langgraph/persistence)
对任务检查点和跨任务记忆的区分,不直接复制其架构规模。
## 运行与维护边界
优先单例、串行、有预算的验证。token/时间限制不能读取订阅余额,不能代替
服务端费用上限。默认不下载大型模型,未知请求结果不自动重试。
原始日志和模型轨迹可能包含私有内容,公开前人工脱敏;见 [安全说明](../SECURITY.md)。
变更通过 PR 检查清单、完整回归和发布证据检查;见 [贡献指南](../CONTRIBUTING.md)。
远程分支保护、私密漏洞报告与账户级设置由维护者确认后配置,文件模板不等于
这些服务端设置已经启用。
+11 -6
View File
@@ -100,7 +100,7 @@ P1 增加模型—工具执行循环、任务状态、预算和恢复能力。
自动测试覆盖主要工程行为;真实网关三次请求均返回 502尚未通过真实模型
端到端验收。失败记录可核查,不能把未返回用量的失败解释为零费用。
## 2026-09-06 follow-up: audit findings and learning workflow
## 2026-09-06 follow-up: installation, execution and verification
Independent, offline reviews reproduced two installation failures: the Python
installer omitted `agent_provider.py`, and skills-folder-only installs could not
@@ -118,7 +118,7 @@ real successful command. Malformed provider batches now stop before dispatch;
valid usage is retained even when response content is invalid. Profile parameters
are either validated and transmitted or rejected, never silently ignored.
`python scripts/run_harness_lab.py` provides a small offline teaching workflow:
`python scripts/run_harness_lab.py` provides a small offline verification workflow:
real missing-asset failure, preparation, a durable pause, a new controller process,
retry and independent verification. Decisions are explicitly scripted, not an
LLM. The final local review run took 1.407 seconds and retained roughly 96 KB including
@@ -128,9 +128,8 @@ directories are refused. No new successful live-model evidence is claimed.
The old MiniSeg preview is retained but labeled as a historical illustration
whose execution provenance was not independently verified. The fixed-commit
repository evidence and media are unchanged. The bilingual
[learning and delivery guide](PROJECT_GUIDE.md) separates current evidence,
planned baseline experiments, model-upgrade gates, career exercises and user
adoption goals. No star-growth or hiring outcome is promised.
[engineering roadmap](ENGINEERING_ROADMAP.md) separates current capabilities,
acceptance boundaries, baseline experiments, model-upgrade gates and maintenance.
The first full regression in this audit passed 60/61 scripts. Its failure
revealed that MSYS2 Python creates `bin/python.exe` despite `os.name == 'nt'`.
@@ -146,5 +145,11 @@ execution evidence, not proof of this new interpreter-isolation assertion.
After these fixes, the complete local suite passed **61/61 scripts in 102.6 s**,
including the actual-venv assertion, installed entrypoints, provider regressions
and teaching lab. `check_publication.py` also verified the unchanged committed
and offline example. `check_publication.py` also verified the unchanged committed
showcase tree. Remote CI for this revision is separate from this local result.
The first remote follow-up (`0b3f4dd`) passed Ubuntu but exposed temporary-path
aliases on Windows/macOS: the folder-install test compared resolved references
to an unresolved boundary, and the offline example compared its Git pointer using
literal path text. Canonical path identity now handles those aliases while
retaining the rule that cleanup removes only this lab's own metadata pointer.
-119
View File
@@ -1,119 +0,0 @@
# RigorPilot: learning and delivery roadmap
[简体中文:完整分析与学习路线](PROJECT_GUIDE.zh-CN.md) · [Engineering record](P0_P1_DELIVERY.md)
Assessment date: 2026-09-06. RigorPilot is best positioned as an **auditable
research-repository execution harness**, not an autonomous scientist. Its useful
differentiator is evidence placed alongside the original README. Its unproven
hypothesis is that this reduces false completion and researcher effort compared
with the same model without the harness.
## Evidence boundaries
Pinned repository snapshots, byte-preserving README annotation, bounded process
execution, recovery and publication checks exist. The four external protocols
include selection-only and partial-training cases; they are not four successful
paper reproductions. Three real provider attempts returned HTTP 502. There is
no successful live-model acceptance or held-out model-quality comparison yet.
The optional transport supports Anthropic Messages; other profile metadata does
not imply another working provider. Local execution is not an OS sandbox.
Ordinary `repro_outputs/ANNOTATED_README.md` output can change the resolution of
upstream relative media paths. The published showcases and lab retain a separate
source-adjacent `RIGORPILOT_README.md`; making that an optional ordinary-run
output, with nested-README/media tests, remains a usability improvement.
## Learn by running
From a clone, with Python 3.11+ and Git:
```bash
python scripts/run_harness_lab.py
```
The lab uses **scripted decisions and real subprocesses**, with no API, GPU or
model downloads. It demonstrates missing-asset failure, preparation, a durable
pause, controller restart and independent verification. Inspect the printed
output's `REPORT.json`, paused/final state, trajectory and source-adjacent
annotated README. Existing output is never overwritten; use
`--output tmp/my-lab-2` for another run. This is an engineering exercise, not
evidence of model intelligence.
Read in order: [skill contract](../skills/ai-research-reproduction/SKILL.md) →
[agent loop](../skills/ai-research-reproduction/scripts/run_agent.py) →
[runtime](../shared/scripts/runtime_runner.py) →
[verifier tests](../scripts/test_agent_runner.py) →
[publication checks](../scripts/check_publication.py).
Predict what happens with an incorrect acceptance condition, write a failing
recovery test yourself, then document a design decision and its trade-offs.
## Delivery gates
| Stage | Scope | Acceptance |
|---|---|---|
| Reliability | Installation, false-success and configuration regressions; lab | Full tests, installed entrypoints and failure cases pass |
| Live canary | One pinned micrograd run, roughly half to one engineering day | Actual model/tool trace, usage, verifier result and source identity |
| Pilot comparison | Six tasks × three conditions, roughly 24 days | All failures retained, reviewed graders, success/false-success/intervention/cost results |
| Reusable evaluation | Twelve frozen tasks over 46 repos, key cases repeated three times, roughly 35 days | Separate development and held-out tasks; paired model-upgrade checks |
| User validation | 510 target users, roughly 12 calendar weeks | Time to first success, unassisted completion, repeat usage and failure categories |
These are planning estimates, not promises. Advance from one canary to three
cases before a matrix. Stop on provider failure instead of cycling model names.
User-set provider billing caps remain necessary: token/time gates cannot monitor
subscription balance. Defer unrestricted edits, large GPU runs, multi-agent
debates and long-term memory until evidence justifies their cost and risk.
## Evaluation protocol to implement next
Compare A: same model with generic task instructions; B: A plus RigorPilot skill;
C: B plus durable execution/recovery/evidence mechanics. Freeze model revision,
tool permissions, approved commands, budgets, source/environment and grader.
Capture raw traces for all arms; lack of a polished report must not count as
business-task failure. Different tool permissions turn this into a product
comparison, not a single-factor harness ablation. Pre-reviewed commands do not
test autonomous target discovery.
Cover normal evaluation, missing assets, exit-zero/incorrect result, interrupted
controller, premature completion and an unauthorized large download. Label
injected faults separately. Existing public cases are development/regression
tasks, not unseen holdouts. Record task/split/commit, harness/prompt/grader hashes,
requested/returned model, parameters, cache/hardware, repeats, interventions,
claimed and verified outcomes, traces, tokens, known/unknown cost and latency.
The current model loop's exit/stdout checks are lightweight execution criteria;
P2 needs task-specific external graders for actual artifacts, metric tolerances
and evaluation conditions, protected from modification by the executing agent.
Report task success and false success with counts; safe and incorrect blocking;
cost per success including failed attempts; duplicate execution after recovery;
and evidence integrity separately. Include provider failures in user-facing
success denominators, with conditional availability metrics alongside them.
Small pilots warrant per-case evidence, not broad generalization percentages.
Freeze primary metrics and stop rules before testing the holdout.
See [Anthropic's evaluation guidance](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents).
## Architecture and model iteration
Borrow testable acceptance contracts and independent evaluation from
[long-running harness practice](https://www.anthropic.com/engineering/harness-design-long-running-apps),
not its agent count. Distinguish per-task checkpoints from cross-task memory as
in [LangGraph persistence](https://docs.langchain.com/oss/python/langgraph/persistence).
For a model upgrade: protocol/usage contract test → live canary → paired holdout
comparison. Change one factor at a time. Remove obsolete scaffolding only after
an ablation shows it no longer helps; preserve prior results.
## Career and community value
Tell three evidence-backed engineering stories: a verifier namespace collision
that falsely passed an unexecuted task; uncertain-dispatch recovery without
blind replay; and local-versus-installed/published artifact discrepancies.
These demonstrate agent application and evaluation engineering, not distributed
production scale or model-training research. The public
[Model Evaluations role](https://job-boards.greenhouse.io/anthropic/jobs/5198255008)
is one useful skills reference, not a hiring guarantee. Explain your own decisions,
tests and AI collaboration honestly; do not invent improvement or adoption metrics.
Keep the two installation commands and real evidence cards. Label the legacy
MiniSeg preview as an unverified historical interface illustration. Publish a
short genuine end-to-end demonstration only once it runs; invite users with
maintainer approval and collect [reproduction feedback](https://github.com/lllllllama/RigorPilot-Skills/issues/new?template=reproduction.yml).
Measure first success and repeat use before optimizing promotion. Stars are an
attention signal, not usage or a guaranteed outcome.
-175
View File
@@ -1,175 +0,0 @@
# RigorPilot学习、验收与产品路线
[English](PROJECT_GUIDE.md) · [首页](../README.zh-CN.md) · [工程记录](P0_P1_DELIVERY.md)
本评估基于 2026-09-06 的代码和公开证据。下面的时间是单人开发估算,
不是完成承诺;真实模型服务、第三方仓库和用户反馈可能改变排期。
## 1. 项目究竟解决什么问题
定位为**可审计的科研仓库执行 Harness**:把 README 中已审核的运行目标,
变成有边界、可恢复、能核查的执行记录。不是通用自主科学家,也不是仅靠
提示词改善回答的包装。最适合小型评测、复现前的可行性检查、训练启动核验。
价值假设是:研究者能更快辨别“运行成功、部分完成、环境阻塞、指标不匹配”,
减少误报与人工排查时间。这个假设仍需要对照实验和外部用户验证。
| 维度 | 已有基础 | 不能据此声称什么 |
|---|---|---|
| 原文与展示 | 固定版本仓库、保留文件和媒体、插入块可剥离并还原 README 字节 | Markdown 字节完整不保证所有上游远程媒体永久在线 |
| 执行机制 | 真实子进程、状态、事件、取消、恢复与命令预算 | 本机执行不是 OS 沙箱;预算检查不是硬磁盘配额 |
| 决策层 | Anthropic Messages 工具循环、已审核命令选择、独立完成检查 | 不能自主修改源码;其他 provider 的 metadata 不等于调用支持 |
| 现有外部评测 | micrograd 测试minGPT 仅选择;两项短训练的部分记录 | 四项协议通过不是四篇论文复现,更不是陌生任务成功率 |
| 真实模型 | 三次 HTTP 502 失败记录保留 | 没有成功的 live model 验收;未返回用量不代表免费 |
| 学习实验 | 可运行的离线失败—恢复实验,真实进程、模拟决策 | 不是模型智能、自动诊断或学习能力证明 |
最强的差异化是“把执行证据放回研究者本来就在看的 README”。最弱的是
“这些护栏究竟比同一个模型裸跑带来多少收益”尚未量化。继续堆 skill、模型名、
截图或多 Agent 数量不能回答后一个问题。
还需注意普通运行的 `repro_outputs/ANNOTATED_README.md` 位于子目录,
上游相对媒体路径可能无法在那里直接显示;现有公开展示和教学实验另外提供
源文件旁的 `RIGORPILOT_README.md`,保留了正确上下文。把这种可选发布能力
推广到普通入口,配合嵌套 README/相对图片测试,是后续首次体验优化项。
## 2. 先亲手跑通,再读架构
从克隆后的项目根目录运行Python 3.11+、Git无 API、GPU、模型下载
```bash
python scripts/run_harness_lab.py
```
这是专门标注的教学夹具,预置决策,真实执行。输出路径和证据入口由命令打印;
已有目录不会覆盖。读 `REPORT.json`、暂停状态、最终 `agent_state.json`
`trajectory.jsonl` 和相邻批注 README解释为什么有三次命令执行但只有一次准备。
`--output tmp/my-lab-2` 创建新的实验,不把模拟结果混入真实模型 benchmark。
```mermaid
flowchart LR
A[README 与已审核任务] --> B[模型或教学决策器]
B --> C[工具权限与预算]
C --> D[持久化进程运行时]
D --> E[观察与状态检查点]
E --> B
E --> F[独立验收]
F --> G[原文增量批注与证据]
```
按顺序阅读,每层只回答一个问题:
| 层 | 入口 | 你要能独立解释 |
|---|---|---|
| Skill 契约 | [SKILL.md](../skills/ai-research-reproduction/SKILL.md) | 谁决定权限?为什么默认不探索? |
| 任务与模型 | [run_agent.py](../skills/ai-research-reproduction/scripts/run_agent.py)、[agent_provider.py](../shared/scripts/agent_provider.py) | 模型能选什么?配置是否真的被发送? |
| 执行与恢复 | [runtime_runner.py](../shared/scripts/runtime_runner.py) | 进程执行了但结果没存下时,能不能重跑? |
| 独立验证 | [test_agent_runner.py](../scripts/test_agent_runner.py) | 为什么模型说“完成”不能直接变成 success |
| 证据与发布 | [check_publication.py](../scripts/check_publication.py) | 为什么本地文件存在不代表 GitHub 用户能看到? |
| 全生命周期 | [run_all_tests.py](../scripts/run_all_tests.py) | 改动如何经过本地回归、安装验收、三平台 CI |
建议你亲自完成三项练习,保留自己的分析而不只保留 AI 生成代码:
1. 在新的教学实验中把验收条件改错,预测结果,再解释为什么不能成功。
2. 为一个尚未覆盖的恢复边界写失败测试,先看测试失败,再改实现。
3. 对一次改动写半页 ADR问题、两个候选方案、取舍、验证、未解决风险。
## 3. 以验收门槛推进,而不是以功能数量推进
| 阶段 | 范围与估算投入 | 完成门槛 |
|---|---|---|
| 现在:可靠性与首次体验 | 修安装、假成功、配置失真;补离线学习实验 | 全套回归和安装路径通过;坏输入不能变成成功;不覆盖用户输出 |
| 下一步:真实模型 canary约半天至一天 | 可用 endpoint 下只跑一例 micrograd | 保存实际模型回复、工具轨迹、退出码、用量、验证结果和原文哈希;无人工替写轨迹 |
| 小型效果对照,约 24 天 | 6 个任务 × 3 条件 × 首轮 1 次 | 所有失败入账;逐项人工核查 grader同时报告正确成功、误报、干预、时间和费用 |
| 回归评测集,约 35 天 | 12 个冻结任务、46 仓库;重要项重复 3 次 | 开发集与保留集分开;新增 bug 进入回归集;模型升级可成对比较 |
| 真实用户试用,约 12 周日历时间 | 510 位目标用户、一个版本、一条短任务 | 记录首次成功率、首次成功时间、重复使用和失败类别;收集同意后才引用反馈 |
优先串行,先 1 例、再 3 例、再小矩阵。首次真实调用仍然 502 时立即记录并停下,
不要反复切模型碰运气。费用预算由用户在服务端设置Harness 的 token/时间上限
不等于订阅余额监控,不能保证“余额不低于 60%”。
暂缓:大规模 GPU 训练、自由源码修补、自动下载大模型、多代理辩论、向量记忆库、
重型可视化平台。缺 OS 隔离时不要向陌生仓库开放任意执行权限。
## 4. 怎样证明 Harness 真的有用
这是待实现的 P2 实验协议,不是已有成绩。先固定任务、资源和验收器,再跑模型。
三条条件A 为同模型+通用任务提示B 为同模型+RigorPilot skill 指令;
C 为 B+持久化执行/恢复/证据机制。使用相同模型版本、工具权限、审核命令集合、
终止预算、初始仓库与环境。A/B 也由实验控制器保留原始轨迹,不能因没有漂亮
报告就被判业务失败。若工具权限或命令集合不同,标记为产品端到端比较,
不得解释为 Harness 单因素收益。命令已预审的任务也不能证明自主目标发现能力。
第一批六种任务:正常小评测、缺资源后准备重试、程序成功但验收不符、
命令成功后控制器中断、模型提前宣告成功、需要未授权的大额下载。
真实仓库任务与注入故障分开标记;精确 commit、测试条件、故障注入和目标结果
在运行前冻结。现有公开四仓库用作开发/回归,不冒充从未看过的保留集。
当前模型循环的退出码/输出字串检查只覆盖轻量执行条件。P2 应增加任务专用的
外部验收器:检查实际产物、指标容差和评测条件,保护评分器不被执行中的 Agent
修改。不要把“打印了成功文字”当作完成了研究任务。
每条结果至少记录task ID、split、repo commit、harness/prompt/grader hash、
模型请求名与实际返回名、参数、依赖/缓存/硬件条件、重复编号、人工干预、
模型自报结论、独立 verdict、事件链接、输入/输出 token、费用是否已知、耗时。
provider 失败单列,但不从所有任务的用户成功率分母中删除;另报服务可用时成功率。
重点指标:
- 正确任务完成率;“未完成却宣告成功”的比例与原始计数。
- 安全阻塞率与错误阻塞率,避免靠全部拒绝刷安全分。
- 每个成功任务的成本(包含失败尝试成本)、人工介入次数、恢复重复执行次数。
- 证据链接/原文完整性另列;运行成功、研究指标匹配、报告完整不是同一个分数。
小样本先给逐例结果和不确定性,不宣传泛化百分比。提前写明主要指标和停止规则;
调 prompt 时只看开发集,冻结后才看保留集。任何误报成功先阻止发布;如果成功率
持平但人工排查时间明显降低,也可以形成有效产品价值,不必强求分数全面领先。
这与 [Anthropic 的 agent eval 方法](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)
强调的完整执行轨迹和明确判定标准一致。
## 5. 借鉴先进 Harness但保留自己的判断
[Anthropic 的长任务实践](https://www.anthropic.com/engineering/harness-design-long-running-apps)
展示了先约定验收条件、再执行和独立评价的思路,也展示了模型进步后删除旧脚手架的可能。
本项目应该借鉴“可验证的契约”和“按实测删减机制”,不是直接照搬三代理架构。
[LangGraph 的持久化设计](https://docs.langchain.com/oss/python/langgraph/persistence)
区分单任务检查点与跨任务记忆;本项目先把前者做可靠,再考虑后者。
每次模型升级:先做协议/用量契约测试→一个真实 canary→同条件保留集成对对照。
比较正确性、误报、成本和人工干预暂不删除旧结果。Prompt、工具、模型版本
一次只变一个因素。若新模型不再需要某条补救提示,先做消融再删除,不永久堆规则。
不要把 capability 标签或某个 gateway 的兼容性推广为全模型支持。
## 6. 求职:展示完整开发过程,而不是只展示成品
更适合讲述为 Agent 应用工程、评测/可靠性工程、科研工具开发项目,暂不包装为
模型训练研究成果或分布式生产平台。公开的 [Model Evaluations 岗位](https://job-boards.greenhouse.io/anthropic/jobs/5198255008)
关注指标定义、可靠执行、异常归因、实验和沟通;这是能力参照,不是入职保证或国内岗位全貌。
准备三段可现场演示的故事:
1. **防止假成功**:命令名称与内部字段冲突为何会让没运行的任务通过;
先复现,再结构化隔离验证命名空间,用反例测试约束修复。
2. **恢复语义**:为什么不能把一次不确定的执行当作“失败后重试”;
展示检查点、run ID、真实进程输出和不重复执行的证据。
3. **交付一致性**:开发目录能跑、安装目录不能跑,或本地媒体存在、远程缺失;
展示安装隔离测试和 Git tree 哈希检查如何覆盖用户真正拿到的内容。
简历可写“实现 README 原文保留、受限工具执行、持久化恢复、独立验收和跨平台回归”;
只有在自己能解释且亲手验证后才写为个人能力。模型提升百分比、节省费用、
生产用户规模只填真实数据AI 协作方式和自己的设计/验证工作如实说明。
## 7. GitHub把关注转化成第一次成功
首页保持两个安装命令、并列真实仓库示例、短效果说明和证据深链。
历史 MiniSeg 界面示意与固定 commit 的实跑证据分开标记;不美化失败。
新增离线实验服务首次体验live canary 服务真实能力验收,两者不能混用。
优先发布一段 6090 秒未经剪辑伪造结果的真实演示:启动任务→看到状态→
打开批注→点击原始日志。然后围绕同一版本邀请 510 位研究/Agent 开发用户试用,
通过 [复现反馈表](https://github.com/lllllllama/RigorPilot-Skills/issues/new?template=reproduction.yml)
收集阻塞,并把真实失败转为测试。是否发帖、邀请用户和创建 release由维护者决定。
Star 是传播信号,不是使用量。更可控的目标是首次成功时间、无需维护者帮助的比例、
一周后重复使用人数和外部贡献。先让真实用户愿意再次使用,再扩大传播;
任何路线都不能保证高 Star 或面试录用。
+11 -11
View File
@@ -1,16 +1,16 @@
# 离线 Harness 学习实
# 离线 Harness 验证示例
这是一个**模拟模型决策、真实执行进程**的教学仓库,不是科研 benchmark
这是一个**模拟模型决策、真实执行进程**的验证夹具,不是科研 benchmark
也不证明某个模型具备自主诊断能力。它不联网、不调用 API、不下载依赖。
只需要 Python 3.10+ 和 Git请在 RigorPilot 项目根目录运行:
只需要 Python 3.11+ 和 Git请在 RigorPilot 项目根目录运行:
```bash
python scripts/run_harness_lab.py --output repro_outputs/harness-lab
```
输出目录必须不存在;再次验请换一个目录,已有文件不会被覆盖。
输出目录必须不存在;再次验请换一个目录,已有文件不会被覆盖。
通常数秒完成,输出小于 1 MiB。运行器会复制本 README 和两个脚本到输出的
`repo/`,原始文件保持不变。这里是本地教学 fixture不是公开科研仓库复现。
`repo/`,原始文件保持不变。这里是本地 fixture不是公开科研仓库复现。
## 1. 先理解任务
@@ -35,9 +35,9 @@ python evaluate.py
没有 `ready.json` 时,程序以非零退出码报告 `missing asset: ready.json`
准备完成后,它检查数据和结果,输出 `verified: sum=6`
这只是确定性教学验收条件,不是模型精度或论文指标。
这只是确定性执行验收条件,不是模型精度或论文指标。
## 4. 按证据学习
## 4. 检查执行证据
一次运行会经历:读取 README → 记录计划 → 评估失败 → 准备资产 → 暂停 →
新 Python 进程恢复 → 再次评估 → 独立验证。
@@ -51,11 +51,11 @@ python evaluate.py
| `repo/repro_outputs/agent_state.json` | 最终计划、尝试历史和独立验收结果 |
| `repo/RIGORPILOT_README.md` | 原文逐块增量批注;原有相对文件链接仍可打开 |
先对照 `run_harness_lab.py` 的模拟器和真正的 `run_agent.py`,解释两者分别
负责什么。再检查“模型声称完成”和“验证器判定成功”为何不是同一件事。
最后在副本中改变预期输出,观察验证失败;不要把修改后的实验混入原证据。
`run_harness_lab.py` 提供固定决策,真正的 `run_agent.py` 负责工具执行和恢复。
模型声明与验证器结果独立记录。修改验收条件时必须使用新输出目录,
不要把不同配置生成的结果混入同一份证据。
## 5. 这个实验能证明什么
## 5. 验证范围
它检验本机的命令调度、失败记录、跨进程恢复、未重复执行、源码完整性和
README 批注链路。它不能证明真实模型推理、陌生仓库泛化、科研结果复现、
+1 -1
View File
@@ -1,4 +1,4 @@
"""A deterministic teaching check, not a scientific model evaluation."""
"""A deterministic execution check, not a scientific model evaluation."""
import json
from pathlib import Path
+41 -13
View File
@@ -1,5 +1,5 @@
#!/usr/bin/env python3
"""Offline teaching lab: simulated decisions, actual execution and process restart."""
"""Offline verification: simulated decisions, actual execution and process restart."""
from __future__ import annotations
import argparse
@@ -7,6 +7,7 @@ import hashlib
import json
import os
import shutil
import stat
import subprocess
import sys
import tempfile
@@ -34,13 +35,13 @@ class SimulatedDecisions:
"""A fixed script, NOT an LLM. Only command execution and verification are real."""
def complete(self, messages, system, tools, max_tokens, timeout):
steps = [
[call("read_file", "教学模拟:先读取原始说明。", path="README.md"),
call("update_plan", "教学模拟:明确最小任务和验收条件。", steps=["评估", "根据失败准备资产", "恢复并验证"])],
[call("run_command", "教学模拟:有意触发一次缺失资产错误。", command_id="evaluate")],
[call("update_plan", "教学模拟:根据预期的缺失资产错误,先准备再重试。", steps=["已观察缺失资产", "准备资产", "暂停后恢复评估"]),
call("run_command", "教学模拟:只调用 README 中已审核的准备命令。", command_id="prepare")],
[call("run_command", "教学模拟:新进程恢复后只重试失败的评估。", command_id="evaluate")],
[call("finish", "教学模拟:提交独立验证,而非凭声明判定成功。", summary="模拟决策结束;以进程和源码验收为准。")],
[call("read_file", "模拟决策:先读取原始说明。", path="README.md"),
call("update_plan", "模拟决策:明确最小任务和验收条件。", steps=["评估", "根据失败准备资产", "恢复并验证"])],
[call("run_command", "模拟决策:有意触发一次缺失资产错误。", command_id="evaluate")],
[call("update_plan", "模拟决策:根据预期的缺失资产错误,先准备再重试。", steps=["已观察缺失资产", "准备资产", "暂停后恢复评估"]),
call("run_command", "模拟决策:只调用 README 中已审核的准备命令。", command_id="prepare")],
[call("run_command", "模拟决策:新进程恢复后只重试失败的评估。", command_id="evaluate")],
[call("finish", "模拟决策:提交独立验证,而非凭声明判定成功。", summary="模拟决策结束;以进程和源码验收为准。")],
]
index = sum(message.get("role") == "assistant" for message in messages)
if index >= len(steps):
@@ -50,7 +51,7 @@ class SimulatedDecisions:
def lab_task() -> dict:
return {"goal": "离线教学模拟:执行已审核命令,记录缺失资产错误,跨进程恢复并独立验证;不代表真实模型能力。",
return {"goal": "离线验证示例:执行已审核命令,记录缺失资产错误,跨进程恢复并独立验证;不代表真实模型能力。",
"language": "zh", "commands": {
"prepare": {"argv": ["python", "prepare.py"], "documented_command": "python prepare.py",
"expected_stdout": "prepared: ready.json", "timeout_seconds": 10},
@@ -73,9 +74,38 @@ def independent_verification_passed(state: dict, required_commands: list[str]) -
and all(command_checks.get(name) is True for name in required_commands))
def remove_temporary_git_pointer(repo: Path, metadata: Path) -> None:
"""Remove only this lab's regular .git pointer, comparing filesystem identity."""
pointer = repo / ".git"
try:
pointer_stat = pointer.lstat()
except FileNotFoundError:
return
if not stat.S_ISREG(pointer_stat.st_mode):
raise RuntimeError(f"Preserved unexpected .git entry (not a regular pointer file): {pointer}")
try:
lines = pointer.read_text(encoding="utf-8").strip().splitlines()
if len(lines) != 1 or not lines[0].startswith("gitdir: "):
raise ValueError("unrecognized gitdir pointer format")
target_text = lines[0][len("gitdir: "):].strip()
if not target_text:
raise ValueError("empty gitdir target")
target = Path(target_text)
if not target.is_absolute():
target = pointer.parent / target
# Git may canonicalize /var to /private/var or Windows short paths to
# long paths. Relative gitdir paths are relative to the pointer's parent.
actual, expected = target.resolve(strict=True), metadata.resolve(strict=True)
if not actual.is_dir() or not expected.is_dir() or not actual.samefile(expected):
raise ValueError("gitdir does not identify this run's temporary metadata")
except (OSError, UnicodeError, ValueError) as exc:
raise RuntimeError(f"Preserved unexpected .git pointer at {pointer}: {exc}") from exc
pointer.unlink()
def worker(output: Path, resume: bool) -> int:
task = json.loads((output / "TASK.json").read_text(encoding="utf-8"))
profile = normalize_model_profile({"adapter_id": "offline-teaching-simulation", "provider": "simulation",
profile = normalize_model_profile({"adapter_id": "offline-verification-simulation", "provider": "simulation",
"model": "simulation-no-model", "endpoint": "simulation://offline", "capabilities": ["tool_calling"],
"metadata": {"simulation": True, "live_model_evidence": False}})
state = run(task, output / "repo", output / "repo/repro_outputs", profile, SimulatedDecisions(),
@@ -119,9 +149,7 @@ def run_lab(output: Path) -> dict:
if phase == "pause":
shutil.copyfile(repo / "repro_outputs/agent_state.json", output / "CHECKPOINT.json")
finally:
pointer = repo / ".git"
if pointer.is_file() and pointer.read_text(encoding="utf-8").strip() == f"gitdir: {metadata.as_posix()}":
pointer.unlink()
remove_temporary_git_pointer(repo, metadata)
state = json.loads((repo / "repro_outputs/agent_state.json").read_text(encoding="utf-8"))
checkpoint = json.loads((output / "CHECKPOINT.json").read_text(encoding="utf-8"))
annotated = rebase_inserted_evidence_links((repo / "repro_outputs/ANNOTATED_README.md").read_bytes(), "repro_outputs/", "train_outputs/")
+55 -1
View File
@@ -7,7 +7,60 @@ import sys
import tempfile
from pathlib import Path
from run_harness_lab import FIXTURE, ROOT, SOURCE_FILES, independent_verification_passed
from run_harness_lab import (FIXTURE, ROOT, SOURCE_FILES, independent_verification_passed,
remove_temporary_git_pointer)
def test_git_pointer_identity(base: Path) -> None:
repo, metadata, unexpected = base / "repo", base / "metadata", base / "unexpected"
repo.mkdir(parents=True)
(metadata / "child").mkdir(parents=True)
unexpected.mkdir()
pointer = repo / ".git"
for spelling in [str(metadata), "../metadata", "../metadata/child/..", str(metadata / "child" / "..")]:
pointer.write_text(f"gitdir: {spelling}\n", encoding="utf-8")
remove_temporary_git_pointer(repo, metadata)
assert not pointer.exists() and metadata.is_dir(), spelling
for contents in [f"gitdir: {unexpected}\n", "gitdir: ../does-not-exist\n", "not a gitdir pointer\n"]:
pointer.write_text(contents, encoding="utf-8")
try:
remove_temporary_git_pointer(repo, metadata)
except RuntimeError as exc:
assert "Preserved unexpected .git pointer" in str(exc)
else:
raise AssertionError("Unexpected Git pointer was not rejected")
assert pointer.read_text(encoding="utf-8") == contents
assert unexpected.is_dir() and metadata.is_dir()
pointer.unlink()
pointer.mkdir()
try:
remove_temporary_git_pointer(repo, metadata)
except RuntimeError as exc:
assert "not a regular pointer file" in str(exc)
else:
raise AssertionError("Unexpected .git directory was not rejected")
assert pointer.is_dir()
pointer.rmdir()
# Symlink creation may require privileges on Windows; when available,
# metadata aliases must work, while a symlink used as .git is never deleted.
alias = base / "metadata-alias"
try:
alias.symlink_to(metadata, target_is_directory=True)
except OSError:
return
pointer.write_text(f"gitdir: {alias}\n", encoding="utf-8")
remove_temporary_git_pointer(repo, metadata)
assert not pointer.exists() and alias.is_dir()
target_pointer = base / "target-pointer"
target_pointer.write_text(f"gitdir: {metadata}\n", encoding="utf-8")
pointer.symlink_to(target_pointer)
try:
remove_temporary_git_pointer(repo, metadata)
except RuntimeError as exc:
assert "not a regular pointer file" in str(exc)
else:
raise AssertionError("Symlink .git was not rejected")
assert pointer.is_symlink() and target_pointer.is_file()
def main() -> int:
@@ -20,6 +73,7 @@ def main() -> int:
assert not independent_verification_passed({"status": "success", "verification": {
"source_unchanged": False, "commands": {"evaluate": True}}}, ["evaluate"])
with tempfile.TemporaryDirectory(prefix="rigorpilot-lab-test-") as temporary:
test_git_pointer_identity(Path(temporary) / "pointer-fixtures")
output = Path(temporary) / "evidence"
command = [sys.executable, str(ROOT / "scripts/run_harness_lab.py"), "--output", str(output)]
environment = {**os.environ, "GIT_DIR": str(Path(temporary) / "must-not-create.git"),
+8 -15
View File
@@ -104,21 +104,14 @@ def main() -> int:
if "not fields the agent must invent on every run" not in campaign_text:
failures.append("research_campaign optional fields are not clearly downgraded from required")
readme = (repo_root / "README.md").read_text(encoding="utf-8")
for phrase in [
"RigorPilot Skills",
"Research-first Agent Skills for Deep Learning Experiments",
"Suggested Research Evidence",
"Lifecycle View",
"agent-operating-principles.md",
]:
if phrase not in readme:
failures.append(f"README.md missing `{phrase}`")
readme_zh = (repo_root / "README.zh-CN.md").read_text(encoding="utf-8")
for phrase in ["RigorPilot Skills", "不只是更高分数", "建议的科研证据体系"]:
if phrase not in readme_zh:
failures.append(f"README.zh-CN.md missing `{phrase}`")
# Product copy and heading names may change. Test discoverable guidance,
# not legacy marketing phrases that force long content back onto the page.
for name in ("README.md", "README.zh-CN.md"):
readme = (repo_root / name).read_text(encoding="utf-8")
for reference in ("references/agent-operating-principles.md",
"references/research-rigor-principles.md"):
if f"]({reference})" not in readme or not (repo_root / reference).is_file():
failures.append(f"{name} must link to existing guidance: {reference}")
print(f"ok: {not failures}")
print(f"public_skills: {len(public_names)}")
+3 -1
View File
@@ -31,7 +31,9 @@ def invoke(script: Path, args: list[str], cwd: Path) -> subprocess.CompletedProc
def main() -> int:
root = Path(__file__).resolve().parents[1]
with tempfile.TemporaryDirectory(prefix="rigorpilot-folder-install-") as temporary:
workspace = Path(temporary)
# Resolve platform aliases (/var vs /private/var, Windows short paths)
# before comparing this boundary with resolved installed references.
workspace = Path(temporary).resolve()
installed = workspace / "agent-home" / "skills"
shutil.copytree(root / "skills", installed,
ignore=shutil.ignore_patterns("__pycache__", "*.pyc"))