* evals: add `claude plugin eval` suite for the academic-paper citation-check flow Eight cases (six fire, two negative) under plugin-evals-citation-check/, graded as a with/without-plugin ablation. Fire cases carry a synthetic source pack so the four author-defined citation failures (no source on file, wrong authors, hedged finding cited as established, retracted or concern-flagged paper cited as live) are detectable offline. Styles: APA 7 (en / zh-TW mixed / es), IEEE, Vancouver (style unnamed), Chicago NB. Cases pin model: sonnet; run with --judge-model opus. Calibration (two pilots, 2026-09-13) and caveats are in the suite README. The with-plugin arm cannot load the mode prompt in the eval sandbox because the command stub references plugin files by relative path (#857), and plain-language prompts fired the skill in 3 of 6 cases (#858; the Spanish case is one data point for #850). No uplift figure is claimed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH * evals(citation-check): apply cross-model review findings and re-calibrate Twelve of thirteen review findings applied: replace the disputable four-author "et al." planting in 04 with a year mismatch; make the clean citations in 02 and 06 supported by their abstracts; replace real Taiwan journal names in 02 with fictional ones; tie the 08 presence regexes to an "unused" statement; require metadata preservation and reject audit content on the 07 conversion negative; add no-overreach to 05; turn the 02 language check into an llm grader; exempt unchanged entries in a complete corrected list from no-false-positive and drop its DOI claim; drop the 04 style-name regex (both arms fixed the year without naming the style). The one rejected finding (08 skill check "display-only") was wrong: it carries arm: both and is scored; README says so. Pilot 3 on the revised suite: $4.65, max 142 s / 7 turns / $0.45 per run; 04 and 08 re-run clean after the last two grader fixes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EEbmXsHyQ3dzYoep74zSmH --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
4.5 KiB
max_turns, timeout_seconds, allowed_tools, model, runs
| max_turns | timeout_seconds | allowed_tools | model | runs | ||||
|---|---|---|---|---|---|---|---|---|
| 20 | 600 |
|
sonnet | 3 |
請幫我檢查這篇稿子的引用。格式是 APA 7 中文版,照台灣的慣例。附上(一)稿子節錄、(二)參考文獻、(三)來源包:我手上真的有的每一篇文獻,照刊出版本抄的題名、作者、年份、出處、摘要。來源包是完整的,不在裡面的就是我沒有那篇。
(一)稿子節錄
自主學習策略與線上課程完成率呈正相關(王大明、李小華,2022),且不同自主學習剖面的學生在課程完成率上差異明顯(Okafor, 2020)。陳志明、林美玲、張建國(2020)以 1,200 名大學生為樣本,發現時間管理策略與課程完成率呈中度正相關。林美玲(2021)進一步證實,自主學習動機能顯著提升線上學習成效,且效果量達中等以上。跨國比較方面,Kim 與 Park(2021)在韓國與美國兩地的研究顯示,同步互動頻率是線上課程滿意度的最強預測因子;本研究的假設一即以此為基礎。黃世杰(2019)的質性研究則指出,學生對平台介面的熟悉度會影響其自我監控行為。
(二)參考文獻
Kim, S., & Park, J. (2021). Synchronous interaction and satisfaction in online courses: A two-country comparison. Journal of Digital Pedagogy, 17(2), 133–152. https://doi.org/10.5555/jdp.2021.1702
Okafor, C. (2020). Self-regulation profiles of first-year online learners. Distance Learning Quarterly, 33(4), 401–419. https://doi.org/10.5555/dlq.2020.3304
王大明、李小華(2022)。自主學習策略與線上課程完成率之關聯。自主學習研究學刊,25(1),1–28。https://doi.org/10.5555/etr.2022.2501
林美玲(2021)。自主學習動機、學習投入與線上學習成效:以北部某大學為例。北區高教探究,14(3),55–82。https://doi.org/10.5555/jhe.2021.1403
陳志明、林美玲、張建國(2020)。大學生時間管理策略與線上課程完成率。時間管理與學習研究,23(2),77–104。https://doi.org/10.5555/cij.2020.2302
黃世杰(2019)。平台熟悉度與自我監控:線上學習者的質性探究。線上學習實務評論,11(4),23–46。https://doi.org/10.5555/jelt.2019.1104
(三)來源包
-
Synchronous interaction and satisfaction in online courses: A two-country comparison — Kim, S., & Park, J. (2021). Journal of Digital Pedagogy, 17(2), 133–152. 摘要:Across 1,860 students in Korea and the United States, synchronous interaction frequency was the strongest predictor of course satisfaction in both samples (Study 1 survey; Study 2 platform log data). 備註:期刊於 2023 年對本文發布 Expression of Concern,指出 Study 2 的平台日誌資料取得程序有疑慮,調查進行中。
-
Self-regulation profiles of first-year online learners — Okafor, C. (2020). Distance Learning Quarterly, 33(4), 401–419. 摘要:Latent profile analysis of 712 first-year online learners yields three self-regulation profiles; the "planful" profile completes courses at nearly twice the rate of the "reactive" profile.
-
自主學習策略與線上課程完成率之關聯 — 王大明、李小華(2022)。自主學習研究學刊,25(1),1–28。摘要:以 948 名大學生為樣本,自主學習策略量表總分與線上課程完成率呈中度正相關(r = .41),其中目標設定分量表的預測力最高。
-
自主學習動機、學習投入與線上學習成效:以北部某大學為例 — 林美玲(2021)。北區高教探究,14(3),55–82。摘要:以北部某大學 326 名學生為樣本,學習投入對線上學習成效有顯著正向影響(β = .38)。自主學習動機對學習成效的直接效果未達顯著(β = .07,p = .21),僅呈微弱正向趨勢;動機透過學習投入的間接效果達顯著。研究限制一節指出,單一學校樣本限制了推論範圍。
-
大學生時間管理策略與線上課程完成率 — 陳志明、林美玲、張建國(2020)。時間管理與學習研究,23(2),77–104。摘要:以 1,200 名大學生為樣本,時間管理策略與課程完成率呈中度正相關(r = .36)。
-
平台熟悉度與自我監控:線上學習者的質性探究 — 黃世杰(2019)。線上學習實務評論,11(4),23–46。摘要:訪談 28 名線上學習者,發現平台介面熟悉度影響學生的自我監控行為與求助時機。