mirror of
https://github.com/Imbad0202/academic-research-skills.git
synced 2026-09-14 13:51:17 +08:00
docs: Gartenberg et al. (2026) fourth human-in-the-loop anchor, volume non-goal, cognitive-surrender note (#833) (#841)
- README.md / README.zh-TW.md motivation: fourth anchor paragraph for the Organization Science AI Task Force editorial "More versus better" (37(3):795-812). Scope stated: one journal, observational, aggregate, proprietary classifier. Cited as design rationale, not as evidence about ARS output. - POSITIONING.md "Rejected mechanisms": volume as an outcome. No batch manuscript generation, no fan-out of one run into several submissions, time-to-draft booked as a resource cost. - shared/collaboration_depth_rubric.md 1.0 -> 1.0.1: related-construct citation on Cognitive Vigilance (uncritical acceptance of AI output; "cognitive surrender" as the editorial cites Shaw & Nave 2026). Scoring, dimensions, and the descriptive-only reporting rule unchanged; a low score stays an observation, not a failure. - CHANGELOG [Unreleased] > Changed. The claim_strength_ladder.md item from #833 is byte-pinned by the revision-claim-drift suite (CLAIM_LADDER_SHA256 and the frozen v2 adjudication rubric); it is held for a separate decision. Claude-Session: https://claude.ai/code/session_01AYAjWg2eBEz3UV7MZn7eFt Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
committed by
GitHub
parent
27e6c9978d
commit
f832c89f60
@@ -16,6 +16,8 @@ All notable changes to this project will be documented in this file.
|
|||||||
|
|
||||||
### Changed
|
### Changed
|
||||||
|
|
||||||
|
- **Gartenberg et al. (2026) joins the human-in-the-loop anchors as the first journal-side evidence, and volume is recorded as a non-goal (#833).** `README.md` and `README.zh-TW.md` gain a fourth motivation paragraph (the *Organization Science* AI Task Force editorial "More versus better", 37(3):795-812; one journal, observational, aggregate; cited as design rationale, not as evidence about ARS). `POSITIONING.md` "Rejected mechanisms" records the volume non-goal: no batch manuscript generation, no fan-out of one run into several submissions, time-to-draft booked as a resource cost. `shared/collaboration_depth_rubric.md` 1.0 → 1.0.1 adds a related-construct citation to the Cognitive Vigilance dimension (uncritical acceptance of AI output; "cognitive surrender" as the editorial cites Shaw & Nave 2026) without changing dimensions, scoring, or the descriptive-only reporting rule. No new effectiveness claim is made, and no number from the editorial is presented as being about ARS.
|
||||||
|
|
||||||
- **Writer and compiler prompts: unsupported factual claims cannot be rescued by hedging, and generic prose quotas become diagnostics (#825).** The citation-density recovery tree in `academic-paper/agents/draft_writer_agent.md` told the writer to rewrite a claim with no usable source "using hedging language" (its CER-chain fallback row said the same, so did the scored writer contract `shared/contracts/writer/full.json` D2, and rule 5 of the M3 temporal iron rule in the writer and both `report_compiler_agent.md` mirrors allowed a bare hedge when the verifying dates were absent); hedging calibrates uncertainty but cannot supply evidence, so an unsupported premise could pass as recovered. All four sites now route missing factual support to a supporting source or attribution, omission, or an explicit `[MATERIAL GAP]` for author review, and an inference or hypothesis must rest on supported premises and be distinguished from an observed finding. The universal prose quotas in the writer prompt, both `report_compiler_agent.md` mirrors, `academic-paper/references/writing_quality_check.md`, the `academic-paper/SKILL.md` anti-pattern rows, and writer contract D6 (80% TEEL as a scored dimension) are rewritten as context-sensitive diagnostics subordinate to author, venue, and discipline requirements — prompts for judgment, never rewrite gates or a pass/fail score (the exact rules are enumerated in the audit correction). Venue word limits, quote/anchor grammar, protected hedges, and revision authority are preserved. Every live consumer of the reference was checked (`academic-paper/SKILL.md`, `deep-research/SKILL.md`, both compiler mirrors, `writing_judgment_framework.md`, `academic_writing_style.md`); versioned records (README version-history entries, the skills' own changelogs) keep their original wording. `audits/harness-retirement-2026-09-model-update.md` gains an in-place post-release correction naming the exact files and rules the September scan missed (and the #823 / #824 / #826 items). A synthetic held-out scenario set for the unsupported / contradicted-claim recovery path is added under `evals/heldout/unsupported_claim_recovery/` with status `NOT_RUN`; no measured quality improvement is claimed.
|
- **Writer and compiler prompts: unsupported factual claims cannot be rescued by hedging, and generic prose quotas become diagnostics (#825).** The citation-density recovery tree in `academic-paper/agents/draft_writer_agent.md` told the writer to rewrite a claim with no usable source "using hedging language" (its CER-chain fallback row said the same, so did the scored writer contract `shared/contracts/writer/full.json` D2, and rule 5 of the M3 temporal iron rule in the writer and both `report_compiler_agent.md` mirrors allowed a bare hedge when the verifying dates were absent); hedging calibrates uncertainty but cannot supply evidence, so an unsupported premise could pass as recovered. All four sites now route missing factual support to a supporting source or attribution, omission, or an explicit `[MATERIAL GAP]` for author review, and an inference or hypothesis must rest on supported premises and be distinguished from an observed finding. The universal prose quotas in the writer prompt, both `report_compiler_agent.md` mirrors, `academic-paper/references/writing_quality_check.md`, the `academic-paper/SKILL.md` anti-pattern rows, and writer contract D6 (80% TEEL as a scored dimension) are rewritten as context-sensitive diagnostics subordinate to author, venue, and discipline requirements — prompts for judgment, never rewrite gates or a pass/fail score (the exact rules are enumerated in the audit correction). Venue word limits, quote/anchor grammar, protected hedges, and revision authority are preserved. Every live consumer of the reference was checked (`academic-paper/SKILL.md`, `deep-research/SKILL.md`, both compiler mirrors, `writing_judgment_framework.md`, `academic_writing_style.md`); versioned records (README version-history entries, the skills' own changelogs) keep their original wording. `audits/harness-retirement-2026-09-model-update.md` gains an in-place post-release correction naming the exact files and rules the September scan missed (and the #823 / #824 / #826 items). A synthetic held-out scenario set for the unsupported / contradicted-claim recovery path is added under `evals/heldout/unsupported_claim_recovery/` with status `NOT_RUN`; no measured quality improvement is claimed.
|
||||||
|
|
||||||
## [3.21.2] - 2026-09-06 — Model currency for Claude Fable 5.1 and GPT-6 Astra, checkpoint decision provenance, and CJK title-matching repairs
|
## [3.21.2] - 2026-09-06 — Model currency for Claude Fable 5.1 and GPT-6 Astra, checkpoint decision provenance, and CJK title-matching repairs
|
||||||
|
|||||||
@@ -32,6 +32,7 @@ These are not "out of scope" footnotes. They are the load-bearing boundary that
|
|||||||
- **Autonomous experiment execution / coding** (Kong §3.3). An LLM that runs experiments or code without scholar oversight. Rejected — and distinct from the shipped Experiment Provenance Intake (#260): ARS may ingest scholar-declared external experiment provenance and check manuscript claims against the declared results, but it must not initiate, run, modify, iterate, or treat tool-executed experiment / code outputs as evidence inside the pipeline.
|
- **Autonomous experiment execution / coding** (Kong §3.3). An LLM that runs experiments or code without scholar oversight. Rejected — and distinct from the shipped Experiment Provenance Intake (#260): ARS may ingest scholar-declared external experiment provenance and check manuscript claims against the declared results, but it must not initiate, run, modify, iterate, or treat tool-executed experiment / code outputs as evidence inside the pipeline.
|
||||||
- **Physical wet-lab automation API** (Kong §7.4.6). An interface that drives liquid handlers or automated labs. Rejected: even with safeguards, this extends beyond a research copilot's scope into laboratory infrastructure, and conflicts with the copilot-not-pilot positioning.
|
- **Physical wet-lab automation API** (Kong §7.4.6). An interface that drives liquid handlers or automated labs. Rejected: even with safeguards, this extends beyond a research copilot's scope into laboratory infrastructure, and conflicts with the copilot-not-pilot positioning.
|
||||||
- **Simulated human-subjects review committee.** LLM lenses named after statutory committee seats, pre-committing a protocol risk level and combining seat judgments into a committee-like result. Rejected: statutory composition rules create an independent, representative, conflict-accountable human body; they are not an epistemic recipe whose legitimacy transfers to model personas ([45 CFR 46.107](https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-A/part-46/subpart-A/section-46.107); [Taiwan Human Subjects Research Act, Art. 7](https://law.moj.gov.tw/ENG/LawClass/LawAll.aspx?pcode=L0020176)). A risk level cannot be meaningfully pre-committed before protocol facts are seen, determination letters are not ethical ground truth, and the unresolved reviewer severity-band error (#648) is especially consequential when risk language is the output. The ownership boundary is categorical: AI may generate questions or advisory observations, but a judgment that binds an absent person requires an accountable human owner. If this topic returns, the defensible object is an RFC and held-out evaluation of multi-lens *question generation*—concern recall, false reassurance, and abstention—not risk levels, committee verdicts, or a system called a committee.
|
- **Simulated human-subjects review committee.** LLM lenses named after statutory committee seats, pre-committing a protocol risk level and combining seat judgments into a committee-like result. Rejected: statutory composition rules create an independent, representative, conflict-accountable human body; they are not an epistemic recipe whose legitimacy transfers to model personas ([45 CFR 46.107](https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-A/part-46/subpart-A/section-46.107); [Taiwan Human Subjects Research Act, Art. 7](https://law.moj.gov.tw/ENG/LawClass/LawAll.aspx?pcode=L0020176)). A risk level cannot be meaningfully pre-committed before protocol facts are seen, determination letters are not ethical ground truth, and the unresolved reviewer severity-band error (#648) is especially consequential when risk language is the output. The ownership boundary is categorical: AI may generate questions or advisory observations, but a judgment that binds an absent person requires an accountable human owner. If this topic returns, the defensible object is an RFC and held-out evaluation of multi-lens *question generation*—concern recall, false reassurance, and abstention—not risk levels, committee verdicts, or a system called a committee.
|
||||||
|
- **Volume as an outcome.** Batch-generating manuscripts, fanning one research run out into several submissions, or treating time-to-first-draft as a result to optimize. Rejected: ARS never batch-generates manuscripts, never drives multiple submissions from one run, and books time-to-draft as a resource cost, not an outcome; every run is one scholar's one manuscript, with the scholar confirming each stage transition. The external reason to say this out loud is journal-side: Gartenberg et al. (2026, *Organization Science* 37(3), [10.1287/orsc.2026.ed.v37.n3](https://doi.org/10.1287/orsc.2026.ed.v37.n3)) read one journal's 2021–2026 submission and review corpus as moving toward "more rather than better" research under current AI tools and publication incentives. That evidence is observational, aggregate, and from a single journal; ARS cites it as rationale for this boundary, not as a claim about its own output.
|
||||||
|
|
||||||
These are first-party scope boundaries and review criteria for future changes, not runtime guarantees. First-party ARS treats each as out of scope; adding one would require changing this recorded boundary, not merely adding a feature.
|
These are first-party scope boundaries and review criteria for future changes, not runtime guarantees. First-party ARS treats each as out of scope; adding one would require changing this recorded boundary, not merely adding a feature.
|
||||||
|
|
||||||
|
|||||||
@@ -34,6 +34,8 @@ v3.8 closes the second half of the L3 gap. v3.7.3 made every citation carry a lo
|
|||||||
|
|
||||||
[**Ren et al.**](https://arxiv.org/abs/2607.13104) (2026, *Self-Improvements in Modern Agentic Systems: A Survey*) supplies a third, survey-level anchor. Its scientific-discovery synthesis (§7.4) concludes that discovery agents cannot easily verify novelty, correctness, or reproducibility on their own and may exploit weak proxies instead, must manage evidence across heterogeneous tools and literature, and raise governance issues — "scientific writing can also amplify misinformation when the evidence is weak." Its generation-loop chapters (§5.1–§5.2) list human auditing and retained human anchors among the practical safeguards for self-generated evaluation loops, and its historical chapter (§2.2) records the oldest form of the same lesson: the practical success of Lenat's EURISKO depended heavily on the user serving as the external evaluation signal, pruning unproductive heuristic drift — a limitation the survey notes persists in modern agentic systems. ARS cites the survey as design rationale for its human-in-the-loop stance, not as empirical proof that human-in-the-loop pipelines outperform autonomous ones; the survey's actionable deltas for ARS are tracked in #539–#541 and #547–#550.
|
[**Ren et al.**](https://arxiv.org/abs/2607.13104) (2026, *Self-Improvements in Modern Agentic Systems: A Survey*) supplies a third, survey-level anchor. Its scientific-discovery synthesis (§7.4) concludes that discovery agents cannot easily verify novelty, correctness, or reproducibility on their own and may exploit weak proxies instead, must manage evidence across heterogeneous tools and literature, and raise governance issues — "scientific writing can also amplify misinformation when the evidence is weak." Its generation-loop chapters (§5.1–§5.2) list human auditing and retained human anchors among the practical safeguards for self-generated evaluation loops, and its historical chapter (§2.2) records the oldest form of the same lesson: the practical success of Lenat's EURISKO depended heavily on the user serving as the external evaluation signal, pruning unproductive heuristic drift — a limitation the survey notes persists in modern agentic systems. ARS cites the survey as design rationale for its human-in-the-loop stance, not as empirical proof that human-in-the-loop pipelines outperform autonomous ones; the survey's actionable deltas for ARS are tracked in #539–#541 and #547–#550.
|
||||||
|
|
||||||
|
[**Gartenberg et al.**](https://doi.org/10.1287/orsc.2026.ed.v37.n3) (2026, *Organization Science* 37(3):795-812, "More versus better") supplies a fourth anchor, and the first from the journal side. The *Organization Science* AI Task Force scored every first submission (6,957) and every text-format review (10,389) the journal received between January 2021 and February 2026 with a commercial AI-writing classifier and standard readability indices. Manuscripts scored as heavily AI-written read worse on those indices and were desk-rejected more often; reviews scored as more AI-written leaned toward theory and away from data; and the editors conclude that current AI tools, amplified by publish-or-perish incentives, "appear to be pushing the system toward an equilibrium of more rather than better research." Their §5 contrasts "cognitive surrender" (Shaw & Nave, 2026, as cited there) with human-first use and asks authors to disclose how a manuscript was produced. The evidence is observational, aggregate, and from one journal, and the classifier is a proprietary instrument. ARS cites the editorial as design rationale for recording volume as a non-goal (see `POSITIONING.md`) and for the Collaboration Depth Observer and the claim-strength ladder, not as evidence about ARS output; the actionable deltas are tracked in #829–#833.
|
||||||
|
|
||||||
v3.3 was inspired by [**PaperOrchestra**](https://arxiv.org/abs/2604.05018) (Song, Song, Pfister & Yoon, 2026, Google): Semantic Scholar API verification, anti-leakage protocol, VLM figure verification, and revision-trajectory tracking. ARS now implements that last idea through categorical, evidence-anchored criterion trajectories rather than score deltas.
|
v3.3 was inspired by [**PaperOrchestra**](https://arxiv.org/abs/2604.05018) (Song, Song, Pfister & Yoon, 2026, Google): Semantic Scholar API verification, anti-leakage protocol, VLM figure verification, and revision-trajectory tracking. ARS now implements that last idea through categorical, evidence-anchored criterion trajectories rather than score deltas.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -34,6 +34,8 @@ v3.8 補上 L3 缺口的另一半。v3.7.3 讓每一筆引用都帶 locator anch
|
|||||||
|
|
||||||
[**Ren 等人**](https://arxiv.org/abs/2607.13104)(2026,*Self-Improvements in Modern Agentic Systems: A Survey*)補上第三個、survey 層級的錨點。其科學發現章節的綜合結論(§7.4)指出:發現型 agent 難以自行驗證 novelty、正確性與可重現性,反而可能鑽弱代理指標的漏洞;證據管理必須跨異質工具與文獻維持;並帶有治理疑慮——「證據薄弱時,科學寫作也會放大錯誤資訊」。其生成迴圈章節(§5.1–§5.2)把人工稽核與保留人類標註列為自生成評估迴圈的實務防護;歷史章節(§2.2)則記下同一課題最早的版本:Lenat 的 EURISKO 的實務成功高度依賴使用者充當外部評估訊號、修剪無效的 heuristic 漂移——survey 明言此限制延續到現代 agentic 系統。ARS 引用這篇 survey 作為 human-in-the-loop 立場的設計依據,而非「人機協作必然勝過全自動」的實證證明;survey 對 ARS 可落地的增量記錄在 #539–#541 與 #547–#550。
|
[**Ren 等人**](https://arxiv.org/abs/2607.13104)(2026,*Self-Improvements in Modern Agentic Systems: A Survey*)補上第三個、survey 層級的錨點。其科學發現章節的綜合結論(§7.4)指出:發現型 agent 難以自行驗證 novelty、正確性與可重現性,反而可能鑽弱代理指標的漏洞;證據管理必須跨異質工具與文獻維持;並帶有治理疑慮——「證據薄弱時,科學寫作也會放大錯誤資訊」。其生成迴圈章節(§5.1–§5.2)把人工稽核與保留人類標註列為自生成評估迴圈的實務防護;歷史章節(§2.2)則記下同一課題最早的版本:Lenat 的 EURISKO 的實務成功高度依賴使用者充當外部評估訊號、修剪無效的 heuristic 漂移——survey 明言此限制延續到現代 agentic 系統。ARS 引用這篇 survey 作為 human-in-the-loop 立場的設計依據,而非「人機協作必然勝過全自動」的實證證明;survey 對 ARS 可落地的增量記錄在 #539–#541 與 #547–#550。
|
||||||
|
|
||||||
|
[**Gartenberg 等人**](https://doi.org/10.1287/orsc.2026.ed.v37.n3)(2026,*Organization Science* 37(3):795-812,*More versus better*)補上第四個錨點,也是第一個來自期刊端的錨點。*Organization Science* 的 AI 工作小組用商用 AI 寫作分類器與標準可讀性指標,量了該刊 2021 年 1 月到 2026 年 2 月收到的全部首次投稿(6,957 篇)與文字型審查意見(10,389 份)。被判為大量由 AI 撰寫的稿件在這些指標上更難讀、也更常被 desk reject;AI 味較重的審查意見偏向理論、遠離資料;編輯群的結論是,現行 AI 工具加上「不發表就出局」的誘因,「看來正把系統推向『更多而非更好』的均衡」。其 §5 把「認知投降」(cognitive surrender,該文引 Shaw & Nave, 2026)與「人先行」的用法對比,並要求作者揭露稿件如何產出。這些證據是觀察性、總體層次、且只來自一本期刊,分類器也是專有工具。ARS 引用這篇社論作為「產量不是目標」(見 `POSITIONING.md`)、Collaboration Depth Observer 與 claim-strength ladder 的設計依據,不是關於 ARS 產出的證據;可落地的增量記錄在 #829–#833。
|
||||||
|
|
||||||
v3.3 的靈感來自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(Song, Song, Pfister & Yoon, 2026, Google):Semantic Scholar API 驗證、反洩漏協議、VLM 圖表驗證、修訂軌跡追蹤。ARS 目前以分類式、證據錨定的準則軌跡實作最後一項,不計算分數差。
|
v3.3 的靈感來自 [**PaperOrchestra**](https://arxiv.org/abs/2604.05018)(Song, Song, Pfister & Yoon, 2026, Google):Semantic Scholar API 驗證、反洩漏協議、VLM 圖表驗證、修訂軌跡追蹤。ARS 目前以分類式、證據錨定的準則軌跡實作最後一項,不計算分數差。
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -1,12 +1,12 @@
|
|||||||
---
|
---
|
||||||
rubric_version: "1.0"
|
rubric_version: "1.0.1"
|
||||||
paper_citation: "Wang, S., & Zhang, H. (2026). Pedagogical partnerships with generative AI in higher education: how dual cognitive pathways paradoxically enable transformative learning. International Journal of Educational Technology in Higher Education, 23:11. DOI: 10.1186/s41239-026-00585-x"
|
paper_citation: "Wang, S., & Zhang, H. (2026). Pedagogical partnerships with generative AI in higher education: how dual cognitive pathways paradoxically enable transformative learning. International Journal of Educational Technology in Higher Education, 23:11. DOI: 10.1186/s41239-026-00585-x"
|
||||||
license: CC-BY-NC 4.0
|
license: CC-BY-NC 4.0
|
||||||
---
|
---
|
||||||
|
|
||||||
# Collaboration Depth Rubric
|
# Collaboration Depth Rubric
|
||||||
|
|
||||||
**Status**: v1.0 (introduced ARS v3.5, 2026-04-21)
|
**Status**: v1.0.1 (introduced ARS v3.5, 2026-04-21; patch 1.0.1 on 2026-09-08 adds a related-construct citation, #833)
|
||||||
**Source**: Wang, S., & Zhang, H. (2026). *IJETHE* 23:11. DOI [10.1186/s41239-026-00585-x](https://doi.org/10.1186/s41239-026-00585-x). Open Access, CC BY 4.0.
|
**Source**: Wang, S., & Zhang, H. (2026). *IJETHE* 23:11. DOI [10.1186/s41239-026-00585-x](https://doi.org/10.1186/s41239-026-00585-x). Open Access, CC BY 4.0.
|
||||||
**Canonical location**: `shared/collaboration_depth_rubric.md` in `Imbad0202/academic-research-skills`. External consumers should reference by stable URL; do not vendor (bump the `rubric_version` field on any modification).
|
**Canonical location**: `shared/collaboration_depth_rubric.md` in `Imbad0202/academic-research-skills`. External consumers should reference by stable URL; do not vendor (bump the `rubric_version` field on any modification).
|
||||||
|
|
||||||
@@ -65,6 +65,8 @@ The rubric is **descriptive, not prescriptive**. It does not gate user progressi
|
|||||||
|
|
||||||
**Empirical anchor**: H2a β = 0.437, p < 0.001, f² = 0.243 — **the single highest-impact path** in Wang & Zhang's model. Also the dimension with highest IPMA importance (0.438) *and* lowest IPMA performance (56.7/100), identifying it as the top priority for pedagogical intervention. If a user scores low here, that is the signal the paper most strongly implicates as correctable.
|
**Empirical anchor**: H2a β = 0.437, p < 0.001, f² = 0.243 — **the single highest-impact path** in Wang & Zhang's model. Also the dimension with highest IPMA importance (0.438) *and* lowest IPMA performance (56.7/100), identifying it as the top priority for pedagogical intervention. If a user scores low here, that is the signal the paper most strongly implicates as correctable.
|
||||||
|
|
||||||
|
**Related construct**: the risk this dimension is meant to make visible is *uncritical acceptance of AI output*. Gartenberg et al. (2026, *Organization Science* 37(3), [10.1287/orsc.2026.ed.v37.n3](https://doi.org/10.1287/orsc.2026.ed.v37.n3), §5.1) discuss that risk under the label "cognitive surrender" (Shaw & Nave, 2026, preprint, as cited there) and contrast it with human-first use. Their editorial is observational, aggregate, and limited to one journal's submission and review corpus, and it reads the combination of AI tools and publication incentives, not this construct alone, as tending toward more rather than better research. This note adds a citation only: the dimension's operationalisation, the scoring, and the rubric's descriptive-only reporting rule are unchanged, and a low score remains an observation, not a failure.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
### Cognitive Reallocation
|
### Cognitive Reallocation
|
||||||
|
|||||||
Reference in New Issue
Block a user