全流程科研辅助,一键完成论文产出
为academic-research-skills打分
给出您宝贵的评分:
手机端可长按上方图片保存到相册,或点击「下载/分享」分享到微信
使用 academic-research-skills,你可以:
Claude Code 专用学术技能,覆盖研究、写作、审稿、修订、定稿全流程,助力高效产出学术论文。
用户评论 (0)
2026年06月16日
2026年06月01日
2026年08月13日
2026年08月03日
2026年08月02日
2026年07月19日
2026年06月17日
v3.21.1
2026年08月24日
What's Changed
- Add OrcaRouter to the community directory — #782
- Update cross-model recommendation surfaces for generation currency — #784
- Repair the Codex ChatGPT-subscription citation transport for codex-cli 0.147.0 — #786
- Record the first Promotion Bakeoff and validate
gpt-5.6-solfor that transport — #788 - Consolidate Markdown parsing helpers and harden link/heading semantics — #791
- Register write-scope guard launcher degradations — #792
- Re-derive
data_access_levelfor standalone paper and reviewer skills — #793 - Add sealed bakeoff and default-off workflow-profile contracts — #795
- Add the opt-in inquiry ledger, alternative-register design freeze, and bounded review-criteria proving set — #796
- Release preparation and version alignment — #797
Highlights
- Repairs the contained ChatGPT-subscription citation transport for Codex CLI 0.147.0, covering stderr authentication attestation, provider-incompatible schema keywords, and the code-mode-host/web-search interaction.
- Adds sealed preregistration for future Promotion Bakeoffs and records the first counterbalanced 180-call bakeoff validating
gpt-5.6-solfor the ChatGPT-subscription citation transport. - Ships a default-off research-workflow profile substrate and the opt-in
ARS_INQUIRY_LEDGER=1alpha, with deterministic contracts, append-only receipts, bounded budgets, fail-visible staleness, and CI-gated conformance. - Freezes the profile-relevant alternative-register design and adds one bounded, source-backed MSR 2027 review-criteria proving set with exact-axis resolution and three-consumer digest binding.
- Consolidates Markdown link/heading grammar, records write-scope guard degradation paths, re-derives data-access annotations, and refreshes cross-model recommendation and community-directory surfaces.
Important boundaries
- The research-workflow profile ships as a default-off, offline-only substrate: no manuscript inference, pipeline hook, or family-specific shipped profile is included, and behavioral usefulness remains
NOT_RUN. - The inquiry ledger is opt-in via
ARS_INQUIRY_LEDGER=1and defaults OFF; its usefulness and usability evidence remainNOT_RUN. - The alternative register is design-only in this release; no runtime implementation ships.
- The review-criteria fixture is one illustrative proving set, not general venue/discipline coverage or a real-author attestation. #575 remains open pending #684's two independent blinded human experts and separate blind adjudication.
- The
gpt-5.6-solresult is transport-qualified: validated for the ChatGPT-subscription citation route only; the first-party API route remains provisional. - No new general research-outcome, safety, reviewer-correctness, or effectiveness claim ships.
Full technical details: CHANGELOG.md
详细ChangeLogv3.21.0
2026年08月18日
What's Changed
- ISO/IEC 42001-spirit gap assessment (dual-track, 2026-08-17) — #762
- Citation-surface version drift + version-consistency invariant 12 (#754) — #763
#743design-doc reset-boundary co-location hotfix — #764- Distribution-surface claims aligned with evidence ceilings (#753) — #766
- Per-channel control-availability matrix (#757) — #768
- DATA_FLOWS.md: single map of network touchpoints + local stores (#758) — #770
academic-pipelinedata_access_levelcorrected toraw+ per-skill pins (#756) — #772- CI workflow enforcement-class table + inventory lint (#755) — #774
- Lightweight risk register + RR-1..3 mirroring lint (#759) — #777
- Solo-maintainer governance statement + SECURITY triage procedure (#760) — #778
- Stage capability/evidence matrix with enforceable claim ceilings (#745) — #751 (+ design freezes #750/#752)
- Pipeline wiring for the #655 claim-standing probe (PR-C) — #733
- Release prep — #779
Highlights
The ISO/IEC 42001-spirit track (epic #761) ships complete. ARS does not pursue certification; it adopts three distilled operating principles — transparency, verifiability, feasibility — with informative anchors to ISO/IEC 42001, and this release closes every finding of the 2026-08-17 dual-track audit:
- Four standing transparency artifacts, each defended by a CI lint:
docs/CONTROL_AVAILABILITY.md(which controls operate in your install channel),docs/DATA_FLOWS.md(what leaves your machine, what is stored, for how long, and how to turn each path off),docs/ARCHITECTURE.md§7.1 (what each CI workflow actually enforces), anddocs/RISK_REGISTER.md(risk → control → evidence status → residual gap, statuses mirrored verbatim from the capability matrix). - Claims aligned with evidence ceilings: distribution manifests drop unlicensed language; the pipeline's no-bypass prose now names its recorded override routes (trust-based controls with audit trails — the five-language README guarantee line included); citation surfaces are lint-pinned to the suite version.
- Governance and security response stated honestly: root
GOVERNANCE.mdrecords decision authority, per-operator cross-model scope (an error-detection control, not organizational independence), release authority, and end-of-life posture, plus the principles↔42001 mapping with Annex C not-applicable assessments;SECURITY.mdbacks its 7-day acknowledgement with a severity-tiered, solo-runnable triage procedure. - Also in this tag: the #745 stage capability/evidence matrix with enforceable claim ceilings, and the #655 claim-standing probe's pipeline wiring (consent-gated, advisory-only, no live provider).
No new effectiveness numbers are claimed in this release; unmeasured surfaces remain explicitly NOT_RUN/DESIGNED in the capability matrix and risk register.
Full changelog: CHANGELOG.md § 3.21.0
详细ChangeLogv3.20.1
2026年08月16日
This patch bundles all release-worthy changes since v3.20.0.
Highlights
- Hardens human-control and integrity contracts across research, revision, and review: an explicit exit from non-generating Socratic RQ mode, reason-bound E6 dispositions, replayable Claim Registry coverage, fail-visible read-scope resolution, criterion-bound categorical reviewer judgements, and replay-valid six-axis review-panel provenance.
- Adds the bounded claim-standing evaluation substrate: consent-bound discovery adapters, versioned stance and evidence contracts, exact-span evidence replay, inert presentation rendering, a bilingual synthetic seed set, and deterministic scoring.
- Adds a closed first-round assignment-ledger gate for the #659 blind bundle, including pair-level exposure blocking, sealed receipts, and isolated write-once delivery.
- Documents an opt-in post-v3.20 roadmap for bounded domain profiles, inquiry branches, alternative registers, and outcome evaluation while preserving the simple default path.
Important boundaries
- These contract changes do not establish improved scientific outcomes, reviewer correctness, complete semantic detection, authenticated human identity, or independent error processes.
- Claim-standing classification remains UNMEASURED: no live stance provider, pipeline hook, relevance assessor, model run, or baseline exists; the discovery adapters have not yet been exercised against live providers. #655 remains open.
- The assignment gate proves structural exposure constraints only; it cannot authenticate that pseudonymous handles represent distinct or independent people. No external sessions, human judgements, model calls, or network runs occurred. #659 remains open.
- The roadmap is planning work, not default-on product behavior; structural expansion remains opt-in and subject to usability and outcome gates.
Full technical details are in the v3.20.1 changelog.
Compare v3.20.0...v3.20.1
v3.20.0
2026年08月14日
v3.20.0 bundles every release-worthy commit merged since v3.19.0. It strengthens evidence, provenance, and authority boundaries across review, revision, citation, human-subjects, and submission workflows while preserving ARS's human-in-the-loop positioning.
Highlights
Evidence-bound review and revision — reviewer and re-review contracts gain role-scoped criteria, typed evidence anchors, evidence-before-persuasion gates, deterministic receipts, complete retry evidence, author-controlled non-ranking revision roadmaps, and source-bound evidence rows. Unified criteria can now follow one review target across formative, internal, and external review without turning rubric conformance into a substitute for human judgment.
Research-integrity and authority substrates — new or hardened layers cover bibliographic integrity, retraction observations, cross-document consistency, content coverage, human-subjects authority/pathway traces, deterministic submission-packet manifests, committee-correspondence concern accounting, and optional cross-run adjudication-activity observability. These remain advisory or procedural support; institutional determinations and author decisions are not delegated to the suite.
Contained transport and ingestion boundaries — the ChatGPT-subscription citation transport has a closed, bounded protocol with no API fallback; EOF-complete draining rejects late protocol activity and reaps the process group. An opt-in process-isolated PDF text/OCR classifier remains a structural advisory, and the offline claim-standing candidate-ledger substrate supplies contracts and a local finalizer for consent-bound retrieval-evidence records without adding a discovery adapter, live probe, stance judge, or measurement.
Hermetic evaluation substrates — frozen synthetic fixtures, closed schemas, dry-run materializers, durable first-stop evidence, and no-call envelopes now cover role topology, ideation diversity, indirect prompt injection, claim standing, review criteria, and tortured-phrase screening. They make study plans reproducible; they do not by themselves establish safety, efficacy, accuracy, or behavioral improvement.
Access and platform work — Chinese-literature resolution, explicit /ars-* command aliases, Pi integration, clinical-reporting guidance, and platform documentation are expanded without changing the suite's human-in-the-loop model.
Versions
- suite /
academic-pipeline: v3.20.0 deep-research: v2.12.0academic-paper: v3.3.0academic-paper-reviewer: v1.11.0
Important limits
Held-out suites that require independent human experts, judges, or adjudicators remain unmeasured until those people complete the frozen protocols. No API fallback, autonomous OCR-to-anchor gate, institutional authorization, or autonomous publication path is introduced.
Full detail for every change is in CHANGELOG.md.
详细ChangeLogv3.19.0
2026年07月22日
Three advisory-or-opt-in integrity layers plus a launcher fix, bundling every release-worthy commit merged since v3.18.0. All new mechanisms preserve the human-in-the-loop positioning: nothing gates by default.
Highlights
Revision-round claim-drift guards (#569 / #570) — closes the epistemic and token halves of the #390 honest-claim residual: the block-anchored patch confined silent-distortion exposure to touched blocks but never checked a touched block's interior (DELEGATE-52, arXiv:2604.15597). Two complementary layers:
- a claim-strength ladder (
is consistent with < is associated with < predicts < contributes to < affects/leads to < causes) whose invariant is "no silent move, either direction, without an authorizing roadmap item" — wired into revision drafting and a new advisory Phase E6; - a deterministic numeric/citation token-conservation checker (
scripts/check_revision_token_conservation.py) as its necessary-but-not-sufficient complement.
Rather than cite an earlier-generation-model study as motivation, the current frontier model's baseline was measured first (evals/heldout/revision_claim_drift/). Mechanism shape credited to Yila-AI/sci-ssci-skills by @MissOrangePeel.
PDF read-integrity preflight (#512) — a three-signal page-count cross-check (scripts/pdf_read_preflight.py) so a truncated or mispaginated PDF read cannot mint an apparently-valid page anchor. FAIL on positive truncation evidence; UNAVAILABLE (explicit-warning advisory) when verification simply cannot run.
read_scope honest-coverage attestation (#513) — an optional declaration on the human-read ledger (full_text / sections / abstract_only / toc_only) that makes the finalizer's citation promotion read-scope-aware; absent means unknown, never backfilled.
Write-scope guard launcher fix (#545) — removes a watchdog pipe-stall that blocked every healthy PreToolUse write-scope-guard call for the full wall-clock bound (~6 s → ~0.15 s), plus flaky-test margin fixes.
Docs (#564) — de-drifted the SETUP Method 4a description-length figure to a durable comparative form.
Notes
academic-pipeline tracks the suite at v3.19.0; the three underlying skill versions (deep-research, academic-paper, academic-paper-reviewer) are unchanged. Full detail per change in CHANGELOG.md.
v3.18.0
2026年07月18日
External motivation: Ren et al. (2026), Self-Improvements in Modern Agentic Systems: A Survey (arXiv:2607.13104). Eight issues (#539–#542, #547–#550) derived from a two-reader (Claude + codex) full read of the survey, shipped via PRs #551/#552/#554/#555/#556/#557/#558/#559 — every mechanism advisory-or-opt-in, preserving ARS's human-in-the-loop positioning. This release also carries one independent feature: the #544 plugin update reminder.
Highlights
- SessionStart update-available reminder (#543 → #544) — plugin installs now get a one-line reminder pointing at
/plugin update academic-research-skillswhen the installed version is behindmain(24 h cache, 3 s network ceiling,ARS_UPDATE_CHECK=0kill switch). Third-party marketplaces default auto-update OFF and surface no behind-signal — this closes that gap, and this v3.18.0 release is the first it will announce. - Scope-conformance advisory (#547) — per-sub-question scope bindings in the RQ Brief; Phase E4 flags claims that silently broaden beyond their inherited scope as
ADV-E4rows, displayed per-row at the MANDATORY integrity checkpoints. Advisory-only, never gates. - Search-bounded novelty claims (#548) — "first study to…" language defaults to the bounded form filled from the documented search strategy (+ new
last_searched_at); Phase E5 classifiesSUPPORTED_WITHIN_SEARCH/UNRESOLVED, never "globally verified"; the bounding qualifier is compression-protected end-to-end (draft → abstract → formatter). - Risk-stratified Stage 2.5 claim verification (#549) — 100% of HIGH-IMPACT claims + a random sentinel replaces the uniform 30% spot-check, extending the #518 reference tiers to claim level.
- Cache staleness + live re-validation (#541) — cache-through wired into the citation-verification gate by default (closing the v3.11 Delta-2 forward-decl), with an age-based staleness advisory (
ARS_CACHE_STALE_ADVISORY_DAYS) and opt-in per-row live re-validation (ARS_CACHE_REVALIDATE=1). - Cross-model reviewer track (#540) — consent-gated: one seat of the fixed five-seat panel runs on the second model family (explicitly NOT the retired 6th reviewer); single-family runs disclose the correlated-error caveat in a new Review Panel Provenance block.
- Re-review judge independence + Judge Record (#539) — Stage 3' verdicts get an independent judgment-specific cross-model pass with a transparent Judge Record (judge identity, Round-1 panel provenance, rubric surfaces, judging budget).
- Pipeline behavior robustness evals (#550) — metamorphic paired eval seed set for the routing/gate layer; also shipped the reviewer skill's missing zh-TW trigger aliases (審查論文/模擬審查/幫我審這篇…).
- Literature anchor (#542) — the survey joins The AI Scientist and Zhao et al. as the third human-in-the-loop anchor, cited as design rationale, not proof.
Full details: CHANGELOG.md
Quality record: 49 codex gpt-5.6-sol xhigh review rounds across the 8 survey feature PRs + the release PR, all converged to 0 P1/P2; 8 independent security scans, 0 findings; full local suite 3421 passed at tag time.
v3.17.0
2026年07月16日
Security
- Tools allowlist for the three top-level plugin agents (#514, implemented in PR #521 by @madtriceps). The three plugin-exposed agents (
synthesis_agent,research_architect_agent,report_compiler_agent; deep-research sources + byte-identicalagents/mirrors, six files) now declaretools: Read, Write, Edit, Grep, Globin frontmatter — no Bash, no WebFetch/WebSearch — so dispatch-time capability is least-privilege even in hook-less installs, complementing the runtime Bucket A Bash deny (scripts/ars_write_scope_guard.py), which keys on agent name and continues unchanged. Retrospective entry: the code merged just after the v3.16.0 tag; documented here per the changelog-covers-merges gate.
Fixed
-
Blind-checkpoint transport moved to the dispatching layer (#523). The #518 blind disagreement checkpoints told their Bucket A primary owners (
research_architect_agent,editorial_synthesizer_agent) to execute the cross-model curl transport themselves — unexecutable under the runtime Bash deny and, for the architect, the #514 dispatch-time allowlist, so on every hook-active run the check at an irreversible decision silently degraded to single-model, indistinguishable from a transient API outage. New Transport ownership (#523) contract inshared/cross_model_verification.md§ Blind Disagreement Checkpoints: the owner commits its structured decision and emits the sanitized cross-model input as a handoff artifact; the dispatching layer (the main session running the skill, orpipeline_orchestrator_agentin pipeline Mode A — neither is Bucket A) executes § API Call Patterns, applies the mechanical enum comparison, and re-invokes the owner only for the divergence rebuttal; the editorial checkpoint's dispatched shape is an explicit, justified exception to the before-the-roadmap ordering (safe because the sprint-contract boundary keeps cross-model drivers out of the roadmap). The rule generalizes to any Bucket A cross-model owner —devils_advocate_reviewer_agent's independent DA critique routes the same way, with every successful response returned to the owner (no mechanical comparison exists for the dispatcher to resolve); non-fenced owners (integrity_verification_agentat the Stage 2.5/4.5 gates, deep-researchdevils_advocate_agent, the main session) execute directly, unchanged. No Bash/WebFetch re-added to any fenced agent (resolution (a); (c) rejected). Converged 0 P1/P2 across first-party security review + two codexgpt-5.6-solxhigh rounds. -
Pipeline prompt-surface contradictions from the #528 Mode-A replay (#529). The two genuine contradictions of the four ambiguities the replay surfaced: (1) the Methodology Blueprint was listed as a Stage 1 deliverable and in the Material Dependency Matrix but omitted from all three Stage 1→2 handoff surfaces — added to
academic-pipeline/SKILL.md,references/pipeline_state_machine.md, andagents/pipeline_orchestrator_agent.md; (2) the post-review coaching trigger read "After Stage 3 or Stage 3' completion, Decision = Minor/Major", but routing sends a Stage 3' Minor directly to Stage 4.5 — the trigger is now split by stage (Stage 3 = Minor/Major; Stage 3' = Major only) and the Coaching Rules exclusion list extended to match. Text-only. Retrospective entry: the code merged before this entry was written; documented here per the changelog-covers-merges gate. -
Stage 5 / Stage 6 boundary semantics — the two under-specified boundaries from the #528 Mode-A replay (#528). New authority section
references/pipeline_state_machine.md§ Stage 5 and Stage 6 Boundary Semantics, mirrored inacademic-pipeline/SKILL.mdandagents/pipeline_orchestrator_agent.md+references/process_summary_protocol.md. Stage 5: "Before finalization: always MANDATORY" now names exactly one checkpoint — the entry gate between Stage 4.5 PASS and the Stage 5 dispatch, carrying the format decisions; the in-stage content confirmation before the final PDF is Stage 5 execution (not a pipeline checkpoint), and the Stage 5 completion checkpoint (Final Paper delivered, before Stage 6) is FULL — never SLIM — but not MANDATORY. Stage 6: the state machine previously ended atStage 5 → ENDwith no Stage 6 at all; it now defines the Stage 5→6 transition, the decline path (Stage 6 is non-mandatory: declining marks itskippedand the pipeline still terminatescompleted), the terminal checkpoint after the Process Record is delivered, and the canonical terminal-acknowledgement vocabulary (finish/end/done/confirm, or an unambiguous natural-language equivalent) whose acceptance sets the pipeline global state tocompleted. All derived from existing text — no architecture change, no checkpoint relaxed. Thirteen codexgpt-5.6-solxhigh review rounds drove the consequential closure across the wider surface set: thestate_tracker_agentcontract gains Stage 6 (stage_id enum, SSOT block, prerequisite rows, terminal/decline action pairs), the FULL checkpoint-type row stops claiming "before finalization" (it collided with the MANDATORY row), Stage 5 execution consumes the entry-gate citation-style decision instead of re-asking, Stage 6 joins the explicitly-skippable list (the skip validator would otherwise reject the pinned decline path), the engagement-tracking SLIM downgrade gains its FULL-checkpoint exception, and the whole-pipeline collaboration-observer pass is re-timed to Stage 6 record compilation (before delivery — "at pipeline completion" could not coexist with completion-after-acknowledgement). Newscripts/check_pipeline_boundary_semantics.pydefrift lock pins all four #528 resolutions across the five surfaces with 66 mutation tests (one adverse-value witness per closed gap), wired into spec-consistency.yml and the unified pytest manifest; because twelve rounds showed sentence-level pins alone cannot converge on prompt surfaces, all five files also carry bibliography_agent-style whole-file sha256 content locks — any byte change fails CI until the pinned hash is updated in the same commit. Closes #528.
Added
- Canonical cross-model handoff envelope + dispatcher consumer contract (#527). The #523 owner→dispatcher→owner transport path was internally coherent but enforced by prose only — no canonical delimiter, no machine-stable schema, no malformed-result mapping, no pinned consumer trigger, so every test could stay green while a dispatcher silently treated a checkpoint owner's handoff as an ordinary deliverable. #527 closes that: one canonical
[CROSS-MODEL-HANDOFF v1]envelope (checkpoint_kind / owner_agent / correlation_id / expected_result / owner_decision-outside-payload / payload) defined inshared/cross_model_verification.md§ Cross-model handoff envelope, withscripts/cross_model_handoff.pyas the NORMATIVE grammar (parse + outcome routing as pure functions) and a deterministic owner→dispatcher→owner fixture suite on a fake transport (scripts/test_cross_model_handoff.py— no external API, no manuscript upload; literal pins guard the module constants against self-referential testing). The three checkpoint owners (research_architect_agent,editorial_synthesizer_agent,devils_advocate_reviewer_agent) emit the envelope with their closed kind/result-shape pair; the Mode-A orchestrator pins the consumer contract (recognition as a transport request never a deliverable; malformed envelope/result →[CROSS-MODEL-ERROR: malformed_*]→ outcomeunavailable, never a fabricated judgment; agreement → mechanical fill with NO owner re-invocation; divergence → re-invoke the original owner with the minimum return context, the dispatcher never authors the rebuttal; DA full-return: every successful response goes back to the owner;ARS_CROSS_MODELunset stays byte-equivalent). Newscripts/check_cross_model_handoff_contract.pypins the contract across all five surfaces (including a prose-enums-follow-the-module invariant) with a per-branch mutation-witness suite; both suites wired into spec-consistency.yml + the unified pytest manifest. Closes #527. - Defrift lock for the #514 tools allowlist (#524). New
scripts/check_tools_allowlist.py+ a 74-test suite (a failing witness per invariant branch), wired into spec-consistency.yml and the unified pytest manifest. YAML is the authority, not a line scan: every semantic decision reads a duplicate-preserving node tree (yaml.compose, which keeps a shadowed duplicate key visible and resolves an alias into shared node identity), the frontmatter fence is a column-0---only (an indented---inside a block scalar can't truncate the block and hide keys below it), and any frontmatter that will not compose to a mapping or uses a merge key (<<) / alias is a fail-closed error. Invariant 1 pins thetools: Read, Write, Edit, Grep, Globvalue on all six #514 surfaces: the node tree must carry exactly onetoolskey whose value normalizes to exactly the canonical five, plus an additive byte-exact raw-line witness (CR-sensitive, so a symmetric LF→CRLF conversion is drift; fires when the verbatim pinned line is absent). This closes the drift scenario where a future PR edits a source+mirror pair symmetrically (re-adding Bash, dropping a tool, or typoing a name) and passes every CI gate green, becausecheck_agents_mirror_sync.pypins only pairwise byte-equality and the runtime guard keys on agent name, never frontmatter; changing the allowlist now requires touching the lint's pinned value in the same commit (standard lock semantics). Invariant 2 reconciles the frontmatter channel against the runtime channel: any agent whosenameis a Bucket A key inscripts/ars_phase_scope_manifest.jsonmust not declare Bash in atools:key in any YAML-legal form — comma string, quoted string, flow/block list, inline comment,Bash(...)permission specifier (BashOutputis a different tool and not flagged) — failing closed on a missing/non-mapping manifest, unparseable Bucket A frontmatter, or an unrecognizedtoolsshape. Twelve rounds of codexgpt-5.6-solxhigh review plus first-party self-probing drove the design from a line-scan first cut to ayaml.composenode-tree authority (the byte-exact witness anchored to the composed key'sstart_mark.index; frontmatter fences found withsplitlines, which recognizes every YAML line break; the wholetoolsvalue folded through Cf-format-char stripping + NFKC BEFORE any split — so every compatibility separator becomes ASCII first: a fullwidth comma,U+FF0C thatsplit(",")would miss, and a fullwidth-paren specifierBash(git:*)U+FF08/U+FF09 that the(split would miss, both reduce to theirBashbase — so an invisible-character or homoglyph re-spelling of a tool name, or of a separator around it, cannot masquerade as a different token), closing a series of YAML-form fail-opens — quoted/flow/escaped Bash, escaped-key duplicates, merge-key injection (plain<<, chained<<: [*a, *b], alias-to-merged-mapping, a merge tag on a non-scalar? !!merge [x]key, and a merge buried in a sequence), a#-in-quoted-key comment-strip bypass, an indented-fence truncation,!!binary/!!strtag tricks, a leading-BOM skip, parser-dependent duplicatename/toolskeys, nested agent files missed by a non-recursive glob (nowrglob), directory symlinksrglobdoes not descend (now fail closed), bare-CR (old-Mac) frontmatter read as absent, aRecursionErroron pathologically deep nesting (now fails closed), and a zero-width/BOM/fullwidth re-spelling ofBashslipping the exact-string membership test (Bash,Bash,Bash—str.strip()leaves Cf format chars, so the token stayed distinct fromBash; now folded, with the byte-witness confirmed to still fire on an invisible-char canonical value), and the two ordering corollaries where a compatibility separator escaped an ASCII split and soBashwas never isolated as its own token — a fullwidth-paren permission specifierBash(git:*)(U+FF08/U+FF09) past the(split, and a fullwidth commaRead,Bash(U+FF0C) past the,split — both fixed by folding the whole value before any split — plus two byte-witness false-positives (atools:line inside adescription:block scalar, and Unicode line breaks NEL/LS/PS that YAML counts butsplit("\n")does not, both fixed by anchoring the witness to the composed key's byte offset); each closure carries a witness. The round-12 pass established the separator-class fold as complete: an exhaustive first-party scan of NFKC-stable alternate separators (ideographic/Arabic commas U+3001/U+060C, division/fraction slashes, semicolons) confirmed none can isolate a bareBashtoken for any consumer — they do not fold to the ASCII,/(the split (or a normalizing consumer) honors, soRead、Bashstays one non-Bashtoken everywhere — plus a scan confirming no codepoint NFKC-decomposes INTOashand no non-Cfcodepoint folds to empty (no token-merging attack); this boundary is pinned by a documenting non-bug test so the separator set is not later over-broadened into false positives. Plus the one-line allowlist mention the #521 review flagged as absent:docs/PERFORMANCE.md/docs/PERFORMANCE.zh-TW.md§ plugin agents and the SessionStart announce script's plugin-agents line. - Machine-readable degradation registry + omission reason-provenance (#511 Part A). New
shared/contracts/degradation_registry.json: a six-row auditable INDEX of every graceful-degradation mechanism (citation resolver outage →unreachable/unresolvable; contamination-signal API degradation → omit-field; VLM absence → skipped/PASS WITH NOTES; submission-package incompleteness →not_checked/exit 3; cross-model absence → warn-and-continue with required disclosure; non-SR compliance warn-cap — the one legitimate severity cap, because compliance already has a native block>warn>info scale). Each row records failure class → emitted degraded state → diagnostic marker → downstream consumer → terminal-policy effect → the per-mechanism authority as a verbatim CONTENT anchor (line numbers forbidden — the issue's own line refs had drifted by implementation time) → the tests/lints/schemas pinning the behavior. The registry indexes, it never re-authors: semantics stay in each authority file, and newscripts/check_degradation_registry.py(31 mutation tests, spec-consistency + pytest-manifest wired) fails CI when any anchor or pin stops resolving. No score caps — rows use native semantics per the issue's explicit rejection. Plus the one live weakness closed:literature_corpus_entry.schema.jsongains optionalcontamination_signal_omissions(closed enumapi_degradedonly; derivable omissions — manual exemption, no-arxiv-id skip — deliberately not recordable), guarded by manual-entry forbid + per-key mutual exclusion withcontamination_signals+ arxiv-requires-id rules, and written on BOTH paths:bibliography_agentat ingest (sha256 baseline updated per the F2 procedure) and the backfill layer — newcontamination_signals.build_signals_with_omissions()(manual checked upstream so aNonefrom the resolver means exactly "degraded";build_signals_objectbecomes an equivalence wrapper) + idempotentrecord/clearhelpers, with both migration tools recording the omission once on a degraded lookup and clearing it when a later run computes the signal (recovery + idempotency + dry-run tests). A degraded lookup is no longer indistinguishable from "never computed". Advisory-only: R-L3-2-C k/k_max counting is unchanged. Design:docs/design/2026-07-15-511-degradation-registry-design.md. Closes #511. - Transport-fixture integration test for the citation gate (#511 Part B). New
scripts/test_transport_fixture_citation_gate.py+ checked-in redacted raw API bodies underscripts/fixtures/transport_bodies/(success / miss / error per resolver, all metadata synthetic —10.5555example-prefix DOI, fictitious arXiv ID). A URL-dispatch fake aturllib.request.urlopen(any unrouted URL fails the test) feeds the bodies through the four ACTUAL client implementations (crossref_client/openalex_client/semantic_scholar_client/arxiv_client) intoverification_gate.verify_citation, pinning the gate's 3-class verdict end-to-end: all-hit →true(ID fast path, one request per resolver), fabricated IDs (404 + real empty-result bodies through the title fallback) →false, total 5xx outage →unresolvableneververified. Closes the gap where the per-client suites pin one client at a time and the citation eval replays already-reducedresolver_outcomesauthored from the same reducer rule — client parsing was never integration-tested. Wired into the unified pytest manifest. Deliberately NOT a product--offlinemode and NOT a replication of the 51-case gold set (scoped out in #511 as inflation). Part A (degradation registry) remains open. - Executable sprint-contract panel checker (#510). New
scripts/check_panel_synthesis.pyrecomputes both v3.6.2 decision layers from the primary artifacts — Layer 1: each reviewer's own scores → declared fired conditions → own## Editorial Decision; Layer 2: the panel scoring matrix → quantifier thresholds → precedence → the synthesizer's declaredfired_conditions:set AND emitted decision — and fails on mismatch (self-consistency gate on LLM output, not a correctness gate). Exit codes are classified by artifact source (1 synthesis-layer → retry synthesizer once; 2 contract/infra → abort; 3 reviewer-report → unusable reviewer ⇒[PANEL-SHRUNK]; precedence 2>3>1), with a--layer1-onlymode for per-reviewer verification at Phase-2 lint time. The §9 expression vocabulary is implemented as a closed grammar (unknown forms, orphan dimension literals, and empty priority scopes all fail closed — no vacuous truth), report/synthesis output grammars are pinned in the five reviewer agents + synthesizer prompts (role line,score:/fired:lines, exactly-once decision line, fenced-code stripped, duplicate sections rejected), and cardinality is guarded against duplicate paths/byte-identical reports/role-set forgery. Ships with protocol §8.1 runtime wiring, the zero-fired accept-grade fallback aligned across all surfaces (derived from the contract's F0 action, never hardcoded), and the majority quantifier corrected from a⌈N/2⌉+1transcription error to simple majority⌊N/2⌋+1(evidence chain in #531 — every concrete threshold in the v3.6.2 design says 3/5, 2/3). Design:docs/design/2026-07-15-510-panel-synthesis-checker-design.md; cross-model design review (gpt-5.6-sol xhigh) drove the grammar pinning, fired-set verification, and exit-code classification.
v3.16.0
2026年07月12日
Added
-
Model tiering: judgment/execution split with two opt-in directions, default untouched (#517). New
ARS_MODEL_TIERINGenv switch and canonicalshared/model_tiering.md, motivated by Lance Martin's "Cost effective harnesses with Fable" (2026-07-10; advisor-checkpoint configs measured ~90% of frontier-solo quality at ~34% of token cost, with delegation paying only when workers absorb enough tokens to offset per-handoff coordination cost). Default (unset): byte-equivalent pre-#517 behavior — every agent staysmodel: inherit(same opt-in philosophy asterminal_policies).economy(frontier-tier session): the 13 execution-type agents dispatch exactly one tier below the session model, floor Opus-class, never Sonnet (academic-prose tolerance is untested; the article's numbers came from ML tuning) —draft_writerexplicitly flagged as the highest-savings / most quality-sensitive downgrade point.quality-boost(below-frontier session): the judgment-type agents dispatched at the Stage 2.5/4.5 integrity gates and the final-review surfaces step up to the frontier tier; nothing is ever downgraded. Both directions have explicit no-op conditions with a one-line announcement; unknown values warn once and behave as unset (fail-open to the safe default). Tiers are relative positions, never hard-pinned model ids (the v3.7.0opuscommand floor retired in the Fable 5 harness pass is the cited precedent). Because many ARS roles execute inline today (no per-role model choice), the mechanism is dispatch-shaped: when a direction applies to a role, the session dispatches it as a subagent pinned to the target tier — inline roles included — and falls open to inline-on-session-model with a one-line announcement where subagent dispatch is impossible;docs/PERFORMANCE.md's (en/zh-TW) "no separate model routing layer" sentence is reconciled with a pointer. The frozen 39-agent classification (26 judgment / 13 execution — the issue header's 25/12 arithmetic corrected, membership unchanged) lives twice on purpose: a machine-readablescripts/model_tiering_manifest.jsonand the canonical doc's table, pinned to each other AND to the*_agent.mdfiles on disk by newscripts/check_model_tiering.py— set equality with a repo-wide stray sweep (a new skill dir can't smuggle unclassified agents), tier-enum + duplicate checks, and EXACT per-(tier, skill) token-set comparison against the doc table (missing/extra/duplicate tokens, per-row counts, duplicate rows all fail; 15 mutation tests; wired into spec-consistency.yml + the local pytest manifest). Prompt-caching guidance (when a direction is active, route repeated same-stage calls to the SAME worker so its cache accumulates) documented in the canonical doc and each of the fourSKILL.mdfiles' compact## Model Tiering (#517, optional)dispatch block — scoped so the unset default stays byte-equivalent, dispatch shapes included. No agent-file edits (the sha256-lockedbibliography_agent.mduntouched), no schema change, no hook. Spec:docs/design/2026-07-12-517-model-tiering-spec.md. -
Cross-model gate hardening: risk-stratified sampling, blind disagreement checkpoints, id-status allowlist, promotion bakeoff (#518). Four upgrades to
shared/cross_model_verification.mdand its consumers, from a 2026-07-11 cross-model consult (gpt-5.6-sol, xhigh). (1) The integrity-gate cross-model sample moves from uniform random 30% (min 5, max 15) to risk stratification across four mutually-exclusive tiers (highest-precedence tier wins, one verification per reference): HIGH-IMPACT references (headline conclusions, numerical claims, causal claims, methods-critical, disputed) verified 100% uncapped at both gates; a 10% RANDOM sample of the remainder at Stage 2.5 (round-up, min 3, max 10); at Stage 4.5, NEW-CHANGED references (behind claims new or changed since 2.5) verified 100% uncapped plus a 10% CONTROL sample of the unchanged remainder replacing RANDOM — verification budget concentrates where the paper's weight rests, and the results table gains a Tier column (integrity_verification_agentupdated in lockstep). (2) The two irreversible checkpoints — research-design freeze (research_architect_agent) and final editorial decision (editorial_synthesizer_agent) — gain optional blind disagreement checks: the primary commits its own decision in the same structured form first (the architect in a new Design-Freeze Checkpoint Audit blueprint section; the synthesizer's is its emitted decision), the cross-model then produces an independent structured decision from the same inputs (never seeing the primary's decision — same anchoring-prevention rule as the integrity samples; the editorial input is the panel'spanel_sizeN usable reviewer cards, never a hardcoded five), differing enum values trigger a targeted rebuttal addressing each cross-model driver against the evidence on file, and divergence escalates to the user — a review trigger, never a vote, never averaged; under a sprint contract the check runs strictly post-Step-3 against the mechanical protocol'seditorial_decisionand its drivers never enter the scoring matrix. (3) The "6th reviewer — Planned" section is retired, not deferred: the consult's counterproductive-conditions list (score averaging, role duplication, findings treated as confirmed defects, majority-vote false confidence, synthesizer context burn) matches ARS's documented anti-patterns one-for-one; the blind checkpoints are the replacement design, and the live mirrors (.claude/CLAUDE.md,shared/raise_framework.md, SETUP feature tables en/zh-TW) drop the "remains planned" claim. (4) The model-detection snippet separates "which provider endpoint" from "is this id known-good": newCROSS_MODEL_ID_STATUS=validated|provisional|unlistedannouncement with an explicit warning for unlisted first-party-prefix ids (gpt-made-upno longer passes silently) — routing itself is byte-identical, an unlisted id still takes the grounded route and never falls through to the ungrounded compatible branch. Plus a § Promotion Bakeoff operationalizing thegpt-5.6-solprovisional→validated criteria: a 30-reference paired same-day run (20 real / 10 fabricated; committed as a versioned, labeled, sha256-recorded fixture before any run counts; 3 repeats with ≥2/3 majority verdict, a 1–1–1 split scored conservatively against the model that produced it) against five non-inferiority thresholds (grounded-search completion, mismatch recall, false-disagreement rate, jq-guard shape stability — a hard requirement, p95 latency), entry-gated byscripts/cross_model_smoke_test.sh, results recorded underaudits/either way — and a deliberate two-step outcome: a full pass makes the idvalidated, while the recommended default flips only with an additionally stated superiority or operational-benefit reason. Spec:docs/design/2026-07-12-518-cross-model-gate-hardening-spec.md. Related: #517 (model tiering) will reference the checkpoint surfaces added here. -
GPT-5.6 Sol listed as provisional cross-model verifier + explicit reasoning-effort control (#515). OpenAI's
gpt-5.6-sol(released 2026-07-08) joins the canonical model table inshared/cross_model_verification.mdas provisional pending ARS validation — endpoint support (Responses API + hostedweb_search), the reasoning-effort enum (none|low|medium|high|xhigh|max, defaultmedium), and pricing (same standard rates as GPT-5.5; premium isreasoning: {mode: "pro"}on the standard slug, NOT a-promodel id) were verified first-party against OpenAI's model page and GPT-5.6 guide, but ARS-specific behavior (grounded-search completion rate, citation-mismatch recall, false-disagreement rate, jq-guard response-shape stability, p95 latency) has no operating history, so GPT-5.5 stays the recommended default. The documented OpenAI Responses call pattern gains an explicit reasoning-effort control via newARS_CROSS_MODEL_REASONING_EFFORT— set, it is passed asreasoning.effortso the run's effort is visible and reproducible; unset, the field is omitted entirely and each model's own provider default applies (forcing one value would silently change behavior for existinggpt-5.5-pro/legacy setups, a codex-review P2) — and both SETUP quick-setup blocks (en/zh-TW, parity-linted) mirror the new example lines. Newscripts/cross_model_smoke_test.sh— a live, manual (not CI; needsOPENAI_API_KEY) promotion gate asserting HTTP 2xx, a completedweb_search_call, a single verdict token, VERIFIED-carries-source, model echo, and effort echo — is the prerequisite for ever flipping the default to Sol. The canonical doc's Chat-Completions-web-search claim was re-verified against OpenAI's current web-search guide and deliberately left unchanged (a cross-model review suggested it was stale; first-party docs confirm it is still accurate). -
WP advisory held-out miss-rate measurement + acceptance set, Part 2 (#501; direction from the PR #468 review thread, @brycewang-stanford). New
evals/heldout/rq_framing_offlist/: a 48-item held-out set (32 shells outside the WP01-WP20 surface forms and the four in-prompt examples — 23 family variants + 9 off-list — plus 16 domain-native hard negatives), generated cross-model (gpt-5.6-sol), shell items regex-filtered (four negatives intentionally carry listed surface substrings as hard-negative material), dual-annotated with documented drops, English-only per the #468 caveat. Scored against the runtime LLM judge (isolatedclaude-sonnet-5sub-agents, verbatim advisory section only, pre-#503 vs post-#503 variants, two post replicates): overall miss rate 0.34-0.38 (above the inherited FNR < 0.30 line), concentrated in decorated compound-title off-list shells (7/9 missed, stable across replicates; judges read generic topical nouns as the exemption's "specific mechanism"), family-variant generalization under the line post-#503 (0.17-0.22), false-fire 0/16 on both variants. Verdict: miss rate HIGH → per #501's decision rule the set is now the acceptance test for any future advisory change (protocol in the set's README; report ataudits/rq-advisory-heldout-measurement-2026-07-11.md). Deliberately outsideevals/gold/(LLM-judged; notarget.entrypoint, labels not reducer-reproducible). Closes #501; the measured off-list gap is tracked as a follow-up design issue. -
Introduction & Title Rhetoric reference (#500; gap surfaced by PR #485, @lorenzo392). New
academic-paper/references/intro_title_rhetoric_guide.md: CARS three-move Introduction guidance (territory / niche / occupation, with a licensed-gap rule, the universal-negative trap, purpose-sentence discipline, and a common-failures table) plus a title-crafting section (anatomy, four title types with claim-level cautions, checklist, weak→stronger worked examples, and a cross-check against the WP06/WP17/WP18 wording-pattern shells). Wired as adraft_writer_agentStep 1 setup checklist item;academic-paper/SKILL.mdFile Structure reference list updated (stale count 20 corrected to the actual 28). The declined PR #485 skill shape (new top-level skill) stays declined; this lands the two genuinely-uncovered content areas as a reference file per the maintainer response there. -
Korean trigger keywords + routing boundary fixtures (#452 PR 1; proposal and Korean boundary phrases by @devCharlotte, who also authored the native-reviewed Korean README in #469). All four
SKILL.mdfiles gain a**한국어**trigger-keyword line plus a conservative Korean subset in the frontmatterdescription— intent-specific compounds only (논문 심사 / 논문 수정 / 초록 작성 / 체계적 문헌고찰 / 연구부터 논문까지 …), deliberately avoiding the broad standalone terms the proposal flagged (연구, 논문, 작성, 검토). The key 수정-vs-심사 disambiguation lands as two new routing smoke-test fixtures (tests/fixtures/issue_133_routing/09_korean_revision_not_review/,10_korean_review_not_revision/) using the proposer's native-authored phrases; all six boundary cases from the issue pass a routing smoke test on the current primary model (6/6, recorded in the PR). No changes to agents, IRON RULEs, integrity protocols, schemas, modes, or output-language behavior. Closes #452 (the Korean README half shipped in #469/#471).
Changed
-
WP advisory exemption sharpening — decorated title-form shells now caught (#505; direction from the #501 Part 2 measurement). The exemption clause in both
socratic_mentor_agent.mdfiles (deep-research + academic-paper) is narrowed: it now requires a named or operationalized specific (an actual instrument/scale name, a named theory/model/dataset/policy instrument, a named site or population, a specified causal pathway — through what mediator/condition/process A relates to B, not merely that it does — or a stated tension between two identified explanations), declares ordinary domain-flavored topic-label pairs swappable, and adds a decorated-compound-title rule (an evocative pre-colon phrase plus a generic "X and Y (in Z)" subtitle gains no specificity from the decoration — the noun-swap test applies to the part after the colon alone). This closes the failure mechanism the #501 Part 2 baseline measured: judges reading generic topical noun pairs as the exemption's "specific mechanism", which rescued 7/9 off-list shells (six decorated titles plus one interrogative). Measured against the held-out acceptance set per its README protocol in two rounds (initial wording, then a cross-model-review-driven refinement — demographic descriptors excluded from "named population", single-topic subtitles covered — re-measured from scratch; 2 replicates each, same judge model): overall miss 0.375/0.344 → 0.094 in all four post-#505 runs, off-list 0.778 → all 9 items fired in at least one final replicate (final rep2: 9/9), false-fire 0/16 preserved in every run (including the four hard negatives carrying listed surface substrings); no shell missed in both final replicates; on-list gold set unaffected (regex detector untouched, fnr=0/fpr=0). All #505 constraints held: WP table unextended, advisory stays non-blocking and surface-phrasing-only, sentinel contract (test_check_rq_framing_patterns.py) unchanged; the new in-prompt example strings were substring-checked against every held-out item (zero hits) so the set stays held out. Measurement JSONevals/heldout/rq_framing_offlist/measurement-2026-07-11-505.json+ reasoning excerpts appended; report ataudits/rq-advisory-505-exemption-sharpening-2026-07-11.md. Closes #505. -
Reviewer calibration protocol notes LLM-as-judge leniency direction (#484 → PR #506, merged).
academic-paper-reviewer/references/calibration_mode_protocol.mdgains a directional-prior subsection under "Failure cases this mode does NOT fix": when the simulated panel's output is read as a pass/fail signal, assume leniency relative to human expert review until your own calibration shows otherwise, anchored to FARS (Tang et al. 2026, arXiv:2606.31651 — automated reviewer mean 5.00 over 165 papers vs 3.23 paper-level mean from 282 human expert reviews over 140 papers; a descriptive ~1.8-point gap, and the automated score functioned only as a relative ranking). The direction is a working prior (heuristic extrapolation from one measured setup, default-until-measured); the magnitude is explicitly non-portable — never a correction factor or threshold change. Docs only; the panel remains advisory infrastructure behind human checkpoints. FARS added to References. -
WP advisory generalization, Part 1 (#501; direction from the PR #468 review thread, @brycewang-stanford). Both
socratic_mentor_agent.mdfiles (deep-research + academic-paper) now state that the WP01-WP20 table is illustrative, not exhaustive, and name the operative judgment: the noun-swap test (phrasing is shell-like when it survives swapping its nouns for any other field's nouns). Off-list shells that clearly survive the swap may fire the advisory at the same high-confidence bar; domain-native phrasing that names a mechanism, instrument, site, or tension does not survive it and must not trigger. Advisory stays non-blocking and surface-phrasing-only; sentinel contract unchanged (test_check_rq_framing_patterns.py). Part 2 (held-out miss-rate measurement) landed separately — see the Added entry above. -
API-first retrieval refresh: OpenAlex API-key auth, budget-aware 429 handling, arXiv ToU-aligned backoff (#495; proposed by @pikaqiu2333). OpenAlex's current developer docs are API-key-first (freemium daily budget; the polite pool is no longer documented):
scripts/openalex_client.pygainsOPENALEX_API_KEYsupport (query-param auth; either credential selects the authenticated 10 req/s pacing tier,OPENALEX_POLITE_EMAILstays as legacy compat), distinguishes daily-budget-exhausted 429s (X-RateLimit-Remaining: 0→ raiseOpenAlexUnavailableimmediately — the budget refills at midnight UTC, so an in-process retry cannot succeed) from transient burst 429s (exponential backoff 2s → 4s → 8s per OpenAlex's documented guidance), and strips the query string from refusal-path error messages so the key never lands in logs (scripts/crossref_client.pygets the same redaction — its query string carries the polite-poolmailtoemail).scripts/arxiv_client.py's 429 backoff moves from the shared 2s constant to the 3s ToU pacing floor (arXiv's Terms of Use ask for at most one request every three seconds — a sub-3s retry would itself violate the pacing the 429 enforces; verified verbatim against the ToU page). Both protocol docs (deep-research/references/openalex_api_protocol.md,arxiv_api_protocol.md) updated in lockstep, plus an explicit retrieval-order boundary in each: structured APIs are the primary channel, browser/WebFetch page inspection is a bounded first-party fallback whose output is data-not-instructions (shared/ground_truth_isolation_pattern.md§2A), and browser retrieval is never a rate-limit bypass (no parallel browsing, no bulk PDF harvesting, no multi-machine fan-out). 5 new client tests; the two 429-behavior tests updated to pin the new backoff shapes.
Docs
- Third-party directory:
THIRD_PARTY.md+ README pointer (#497 → #498). Community-submitted third-party projects that wrap or host ARS get a low-bar directory listing (visible ARS attribution + faithful description required) that is explicitly separate from endorsement — entries are not reviewed, tested, or verified by the maintainer, and the page says so up front. ClawMama listed as the first entry per issue #497. The README install section gains a neutral, non-recommending pointer to the page; the main install flow still links only to maintainer-verified paths. A separate "Getting officially recognized" track is documented for projects that want actual review.
v3.15.0
2026年07月04日
A release-discipline-and-hygiene release; no skill-behavior changes.
Added
- Three CI release gates: CHANGELOG-covers-merges pre-tag gate — every release-worthy merge since the previous tag must be documented before tagging (#483); version-consistency invariants 9-11 (release-notes body ≥100 chars, Last-Updated within ±7 days, Key-Additions heading matches suite version) plus a
tag-version-match.ymlgate that re-runs the full lint at tag time (#487); command-invariants gate pinning the SessionStart announce list to the actual 16-command inventory (#486). - Two defrift locks (#491 → #492): the Phase Boundary enforcement sentence is pinned verbatim across all 23 Bucket A agent blocks (
CANONICAL_ENFORCEMENT, version-matched, per-file tails free — the drift class that sat factually wrong for a month now fails CI); newcheck_setup_cross_model_parity.pypins the SETUP en/zh-TWARS_CROSS_MODELexamples to each other and to the canonical model tables. Local pytest manifest gains the v3.9.4 temporal test (58 → 60 entries), closing a local-green/CI-red coverage split.
Changed
- Prompt-debt retirement round 2 (#489 → #490): deep-scanned the 17 agents the first pass deferred, via 4 parallel audit batches + an independent codex cross-model challenge. 13 findings (2 P1 + 11 P2; 2 user-rejected and recorded). Both
socratic_mentoragents carried live self-contradictions — stale "quit after 15 rounds" rules against a documented typical 20-30-round run — now merged to a single auto-end authority per file (threshold 30). The repo-wide stale "hook deferred to #134" enforcement sentence (false since PR #294) rewritten at 29 surfaces; few-shot and duplicated-process scaffolds trimmed across 7 agents. The 2026-06-10 F-007 negative-framing deferred item closes as verified-no-rewrite-needed. Audit report:audits/harness-retirement-2026-07-04.md.
Fixed
- DOI badge served from shields.io; link target stays the concept DOI (#482).
- SessionStart announce updated to the full 16-command set.
academic-pipeline tracks the suite at v3.15.0; the other three skill versions are unchanged. Full details in CHANGELOG.md.
v3.14.0
2026年07月02日
What's Changed
- docs: recommend auto permission mode over Skip Permissions by @Imbad0202 in #464
- docs: add GitHub Copilot repository instructions by @Imbad0202 in #465
- docs: add native-reviewed Korean README by @devCharlotte in #469
- docs: credit devCharlotte for Korean README translation by @Imbad0202 in #471
- ci: add platform-port reminder (remind, don't block) by @Imbad0202 in #473
- chore(audit): harness-retirement 2026-07 report (#476) by @Imbad0202 in #477
- chore(prompts): retire expired writing-harness scaffolds in 4 Bucket A agents (#476 P2) by @Imbad0202 in #478
- ci(eval-harness): render PR comment as verdict + table, fold raw JSON by @Imbad0202 in #479
- fix(plugin): declare explicit skill paths in marketplace.json by @Imbad0202 in #480
- docs(release): align all doc surfaces for v3.14.0 by @Imbad0202 in #481
New Contributors
- @devCharlotte made their first contribution in #469
Full Changelog: v3.13.0...v3.14.0
Highlights
- Claude Science importability (#480) — the marketplace manifest now declares explicit skill paths, so Claude Science's "Import from GitHub" finds all four skills (previously zero: GitHub-API importers cannot traverse the symlinked
skills/directory). Verified end-to-end on Claude Science; import guide in README +docs/SETUP.mdMethod 5. Claude Code installs are unaffected. - Readable eval-harness PR comments (#479) — a one-line verdict + per-task table with the raw JSON folded into
<details>, replacing the raw report dump. Display layer only;run_evals, the threshold gate, and the ack contract are byte-identical. - Prompt-debt retirement (#477/#478) — expired writing-harness scaffolds removed from four writer-surface agents (net −111 prompt lines) after the 2026-07 harness-retirement audit, with three-track verification (sub-agent audit + cross-model review + eval harness at 100%).
- Changelog backlog rollup — 16
[Unreleased]entries whose code shipped before the v3.13.0 tag (diff/patch revision mode #390, submission-package verifier #394, eval gold sets #215/#216, and more) are now versioned inCHANGELOG.mdunder a provenance note.