refactor: 删除旧两轴版 cs-onboard,待按 LITE 重做

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
liuzhengdong
2026-06-28 16:17:01 +08:00
parent 5921cb7941
commit 12440e69cd
8 changed files with 0 additions and 1020 deletions
-74
View File
@@ -1,74 +0,0 @@
---
name: cs-onboard
description: 把仓库接入 CodeStable——搭骨架 + 分发共享资产 + 选变更轴载体(GitHub / 本地)+ 归旧档。两条路径自动判断:空仓库从零搭,已有文档走审计 + 迁移。触发:"在这个项目用 CodeStable"、"初始化 / 迁移 CodeStable"。
---
# cs-onboard
把仓库接入 CodeStable 工作流:**搭骨架、分发共享资产、选载体、归旧档**。骨架好了子技能即可运行。本技能只搭环境,不开始干活。
## 启动先扫一次自动判断
不让用户选路径(TA 多半不知道现有哪些文档)。扫:
1. `.codestable/` 在不在 → 不在 = 空仓库;在但不全 = 迁移补齐
2. 旧版目录(`easysdd` / 旧 codestable)→ 提示 `git mv` 到 `.codestable/`,结构兼容
3. Glob 全仓 `.md`(排除 `node_modules/` `.git/`)找零散 spec(`DESIGN/ARCHITECTURE/SPEC/README`、`docs/`)
4. 汇报扫描结论 + 走哪条路径 + 不确定项
迁移详步(审计表 / 逐条对齐 / 处理已有 `.codestable/`)见 `reference.md`。
## 选变更轴载体(onboard 必问一次)
变更轴(issue / epic)落在哪由用户定,写进 `attention.md` 顶部 `载体` 字段:
- `github`:issue / epic 用 GitHub 原生实体;需 `gh` 可用 + 仓库有 remote。**不建**本地 `issues/ epics/`
- `local`:建 `.codestable/{issues,epics}/`,用 md spec + frontmatter status
cs-issue / cs-epic / cs-audit 启动读这字段决定产出形态。
## 标准骨架
```
.codestable/
├── attention.md 启动必读 + 载体字段(最小骨架,实质内容用户后续 cs-note 补)
├── convention.md 体系共识(从技能包分发,勿手改)
├── context/ 现状轴(CONTEXT.md / 取舍说明由 cs-context lazy 建)
├── compound/ 沉淀(cs-keep)
├── issues/ epics/ 仅 local 载体模式建
├── tools/ 共享脚本(分发)
└── reference/ 共享参考(分发)
```
`context/` 下的 CONTEXT.md / 取舍说明不在骨架里——交给 `cs-context` 首次需要时 lazy 创建。
## 分发共享资产(onboard 唯一强制覆盖)
`convention.md` / `tools/` / `reference/` 是技能包维护的资产,权威源在 `cs-onboard/`,项目里只是副本。**一律 shell 整体覆盖,不要 Read+Write**(会截断 / 改缩进 / 费 token):
```bash
cp -f <pkg>/cs-onboard/reference/convention.md .codestable/convention.md
cp -rf <pkg>/cs-onboard/tools/. .codestable/tools/
cp -rf <pkg>/cs-onboard/reference/. .codestable/reference/
```
覆盖前汇报被覆盖文件;用户明说"我改过请保留"才例外。其他已有文件遵守"不经确认不动"。`<pkg>` 一般是 skill 安装目录,不确定先 `ls` 定位;拷完 `ls` 验证。
## 退出条件
通用骨架见 convention,另加:
- [ ] `context / compound / tools / reference` 存在;local 载体则 `issues / epics` 也在
- [ ] `attention.md` 已建且含 `载体` 字段
- [ ] `convention.md` / `tools/` / `reference/` 已从技能包分发
- [ ] 迁移:每条映射有处理结果,没未经确认就动的文件
- [ ] 验收汇报已给
## 容易踩的坑
- 未确认就移动 / 删除已有文件——迁移核心是用户拍板
- 替用户填 attention.md 实质内容——只给模板
- 重新引入 `AGENTS.md` / `CLAUDE.md` 入口——固定 `.codestable/attention.md`
- 建完骨架立刻开干——onboard 是搭环境
- 共享资产走"不覆盖"保守策略——这三样**必须**覆盖,否则停留旧口径
- Glob 忘了排除 `node_modules/` `.git/`
-77
View File
@@ -1,77 +0,0 @@
# onboard 参考模板与迁移详步
## `.codestable/attention.md` 最小模板
attention.md 是 CodeStable 技能启动必读的项目注意事项入口。onboard 创建最小骨架,不替项目 owner 填实质内容;后续短规则由 `cs-note` 追加。顶部 `载体` 字段在 onboard 时确定。
```markdown
# Attention
本文件是 CodeStable 技能启动必读的项目注意事项入口。所有 CodeStable 子技能开始工作前必须读取它。
载体: github | local <!-- 变更轴(issue/epic)落在 GitHub 还是 .codestable 本地 -->
## 项目碎片知识
<!-- cs-note managed: 用 cs-note 维护,新条目按下面分节追加 -->
### 编译与构建
### 运行与本地起服务
### 测试
### 命令与脚本陷阱
### 路径与目录约定
### 环境变量与凭证
### 其他
```
---
## 迁移路径详步
### 步骤 1:生成审计报告
| 现有文件 | 推测内容类型 | 建议归入 | 置信度 |
|---|---|---|---|
| `docs/glossary.md` | 领域术语 | `context/CONTEXT.md`(cs-context 写) | 高 |
| `docs/adr-*.md` | 架构决策 | `context/` 对应取舍说明的「为什么需要灵活性」节 | 中 |
| `docs/feature-*.md` | 功能设计稿 | 已实现 → `context/` 取舍说明;未实现 → 开 issue | 中 |
| `SPEC.md` | 需求? | 需用户确认 | 低 |
置信度:高 = 语义明确匹配;中 = 可推断有歧义;低 = 不明确或多个位置都合理。
### 步骤 2:逐条对齐
中 / 低置信度用 `AskUserQuestion` 问:中 = 给推断理由问"按这个归位?";低 = 描述内容给 2-3 候选 + "跳过"。高置信度不逐条问但在汇报里列,给复审机会。
### 步骤 3:处理已部分存在的 `.codestable/`
- 命名不符规范但有内容 → 提示是否重命名
- 空占位(`.gitkeep` / 空 `.md`)→ 直接补齐不问
### 步骤 4:补齐缺失骨架
对照标准骨架补齐**用户确认后仍缺失**的目录 / 文件,已有内容不覆盖。`convention.md` / `tools/` / `reference/` 例外——这三样无条件用技能包新版覆盖(命令见 SKILL.md)。
### 步骤 5:处理不迁移的文件
用户选"跳过"的:**不移动 / 不删除 / 不重命名**,汇报标"保留原位(未纳入 CodeStable)"。绝不允许未经确认就动。
### 步骤 6:验收汇报
列:迁移清单(from → to)、新建骨架、未迁移文件(保留原位)、变更轴载体、下一步建议。
---
## 旧模型迁移提示
从旧版(feature / roadmap / refactor / requirements 四目录)迁来的项目:
- `requirements/` → `context/`:能力文档剥掉产品口吻重铸成取舍说明;CONTEXT.md / adrs 的术语与决策理由并入 context
- `roadmap/` → 未完成的当 epic 重开;已完成的契约毕业进 context
- `features/ refactors/ issues/` 旧 spec → 未完成的当 issue 重开;已完成的取舍进 context,spec 本身进 git 历史
-64
View File
@@ -1,64 +0,0 @@
# CodeStable 共识与约定
所有 cs-* 技能共享的口径。由 `cs-onboard` 分发到项目 `.codestable/convention.md`,各技能开头一行「遵循 `.codestable/convention.md`」即可,不再各自重复样板。
源真相在技能包 `cs-onboard/reference/convention.md`,维护入口是 `cs-convention`。改了它,已 onboard 的项目重跑 `cs-onboard` 同步。
## 两轴模型
CodeStable 把开发活动建模成两根正交的轴:
- **变更轴**——要做、做完会关闭的事(issue / epic)。是增量。
- **现状轴**——现在是什么、为什么这样(context:词汇表 + 取舍说明)。是这些增量积分出的当前真相。
变更轴关闭时,把毕业的取舍回写现状轴。两轴谁都不讲对方的事:**context 不记历史叙事,issue 不长期描述现状。**
人在环:AI 是高效执行体,程序员对整体把控负责。
## 启动必读
任何 cs-* 技能动手前先读 `.codestable/attention.md`(项目硬约束 + 启动注意 + **变更轴载体**字段)。缺失视为骨架不完整,提示先补齐或跑 `cs-onboard`,不要回退到外部 AI 入口文件。
## 变更轴载体
`attention.md` 顶部记录本项目的变更轴落在哪:
- `载体: github` —— issue / epic 是 GitHub 原生实体,用 `gh` 操作
- `载体: local` —— issue / epic 是 `.codestable/{issues,epics}/` 下的 md spec(frontmatter 带 status)
cs-issue / cs-epic / cs-audit 启动读它决定产出形态。
## 路径与命名
- 所有本地产物在 `.codestable/` 下。
- slug:与当前对话语言一致,一眼看出是什么。英文用小写连字符 kebab-case;中文等直接用词,避免空格和 `/ \ : * ? " < > |`。例外:对外发布的 `docs/` 默认英文(URL / 跨语言协作友好),项目可覆盖。
- 日期:取事情发生 / 提报当天,定了不动。
- 会产生草稿的工作给子目录:主文档对外口径,旁边 `drafts/` 随便堆。
## 单目标规则
每次只动一份文档 / 一件事。一次吐多份用户 review 不过来,最后要么粗糙合入要么放着不看。一次扔多个目标 → 选一个,其余下次。
## 不发散 / 顺手发现
只动该动的。范围外发现值得改的别顺手做,记一条:
> 顺手发现:{位置} {问题简述}。不在本次范围,留后续。
混进来的顺手改让 review 和 git blame 分不清这次到底改了什么。
## 人在环 checkpoint
多阶段流程每阶段末留 checkpoint 让用户把关;拍板(选方案 / 定优先级 / 填没说清的角落)归用户,AI 不自己挑一个掩盖分歧。
## 收尾提交(scoped-commit)
一件事走完把产物提交为一个 commit:范围 = 本次代码 + 相关 spec + 本次实际改过的 context。无关的顺手改不进。提交前用户没明确同意不要 `git commit`。message 一句话说清做了什么。
## 退出条件通用骨架
每个技能退出条件都含(各技能再加自己独有的):
- [ ] 锁定单一目标
- [ ] 用户 review 通过
- [ ] 没顺手改代码 / 其他 spec / 范围外文档
-26
View File
@@ -1,26 +0,0 @@
# CodeStable 维护者说明
由 `cs-onboard` 复制到项目 `.codestable/reference/`。维护技能家族时反复查阅、不适合放各子技能正文的说明。
## 1. 断点恢复
AI 对话随时可能中断(token 超限、网络断开、换设备)。各技能发现不是从零开始时,先检查已有产物完成度,从上次停下处续:
- **context**:某篇取舍说明 / 词汇表已有部分 → 逐节补缺,不重写已完成节
- **issue / epic**:载体上已有记录(GitHub issue 正文 / 本地 md)→ 读完从未完成步骤续;代码已改但收尾记录没写 → 直接补验证 + 记录
- **clarify**:已有 `clarify.md` → 读完问"接着聊还是推翻"
恢复时先简短汇报:"检测到上次到 X 阶段,我从 Y 继续"。
## 2. 扩展点
- **新增子技能**:定型后在 `system-overview.md` 的两轴 / 横切清单 + `cs` 路由表加索引,登记目录位置
- **跨技能新约束**:适用所有技能的规则 → 走 `cs-convention` 写进 `convention.md`,不只改一个技能
- **共享术语**:体系自己形成的稳定术语 → 沉淀进 `convention.md`,别散落重复定义
## 3. 维护规则
- 改体系共享口径走 `cs-convention` 改 `convention.md`,已 onboard 项目重跑 `cs-onboard` 同步副本
- 每次扩展同步更新 `system-overview.md` 索引 + `cs` 路由表 + 相关子技能表述(CLAUDE.md 要求)
- 共享说明优先进 `convention` / `reference`,不散落各子技能
- 每个 `SKILL.md` < 100 行;超了把模板 / 详表拆到同目录 `reference.md`
-44
View File
@@ -1,44 +0,0 @@
# CodeStable 体系总览
CodeStable 是面向严肃工程的 AI 编码工作流。它编排的是**软件本身的生命周期**——不是编排 Agent。人在环:程序员对整体把控负责,AI 是高效执行体。
## 两根正交的轴
开发活动归到两根轴,产物聚在 `.codestable/`:
### 变更轴——要做、做完会关闭的事
增量。落在 GitHub 或本地(onboard 时选,见 `attention.md` 的 `载体` 字段)。
- `cs-issue` — 一件可关闭的变更:bug / 重构 / 小功能 / 杂务,tag 分类型。闭环:记清楚 → 定位 → 改 + 验证 → 关闭回写
- `cs-epic` — 大到塞不进单条 issue 的变更:先定架构(模块拆分 + 接口契约),再拆成带依赖 DAG 的子 issue
- `cs-audit` — 主动扫描发现器 + 对账 context,产出 triage 清单,选中的升级成 issue
### 现状轴——现在是什么、为什么这样
变更积分出的当前真相,落在 `context/`。
- `cs-context` — 领域词汇表 + 取舍说明(happy path / 边界 / 为什么需要灵活性)。只承载代码读不出来的东西,不引用代码位置,不记历史叙事
**两轴的接口**:变更轴关闭时,把毕业的取舍回写现状轴。context 不记历史(那在关闭的 issue 里),issue 不长期描述现状。
## 横切与周边
- `cs-code` — 写代码的纪律(只写当前要的、漂移那刻停)。正交于两轴,任何动手写代码都用
- `cs-keep` — 坑点 / 技巧 / 选型 / 调研沉淀到 `compound/`,纯 markdown,全文检索
- `cs-note` — 一两行启动必读追加到 `attention.md`
- `cs-clarify` — 想法还模糊时的讨论 + 分诊入口,聊清楚后路由到直接写或 `cs-epic`
- `cs-convention` — 维护体系共享口径(分发成 `.codestable/convention.md`)
- `cs-onboard` — 把仓库接入体系
- `cs-doc-tutorial` / `cs-doc-api` — 写给外部读者的指南 / API 参考
## 路由
没有 `.codestable/` → 先 `cs-onboard`。其余诉求由 `cs` 根入口按两轴分诊,详见各子技能。
## 进一步参考
- `.codestable/convention.md` — 体系共识与约定(两轴 / 启动必读 / 载体 / 命名 / 单目标 / scoped-commit / 退出骨架)
- `.codestable/reference/tools.md` — `search-yaml.py` / `validate-yaml.py` 用法(本地载体的 yaml frontmatter;compound 走全文检索)
- `.codestable/reference/maintainer-notes.md` — 断点恢复、新增子工作流登记
- `.codestable/attention.md` — 启动必读项目注意 + 变更轴载体
-82
View File
@@ -1,82 +0,0 @@
# CodeStable 工具用法参考
本文件由 `cs-onboard` 复制到项目的 `.codestable/reference/tools.md`,所有 CodeStable 子技能用项目相对路径 `.codestable/reference/tools.md` 引用。
`.codestable/tools/` 下共享脚本的完整用法参考。子技能里只写本技能特有的 1-2 行典型查询;完整语法和示例看这里。
---
## 1. search-yaml.py
通用 YAML frontmatter 搜索工具。从项目根目录运行,无需安装额外依赖(PyYAML 可选,有则用,无则内建 fallback parser)。
### 基本语法
```bash
python .codestable/tools/search-yaml.py --dir {目录} [--filter key=value]... [--query "全文关键词"] [--sort-by FIELD [--order asc|desc]] [--full] [--json]
```
### filter 语法
- `key=value`:字段精确匹配(大小写不敏感)
- `key~=value`:字符串字段子串匹配;列表字段元素包含匹配
- `key=a|b|c` / `key~=a|b|c`:同一字段多个候选值,候选之间是 OR;在 PowerShell / Bash 中请给整个 filter 加引号,例如 `--filter "status=approved|draft"`
### 排序语法
- `--sort-by FIELD`:按 frontmatter 字段排序(典型字段:`last_reviewed`、`date`、`updated_at`)
- `--order desc|asc`:`desc` 默认,新的在前;`asc` 老的在前(查"谁最久没更新"用这个)
- 字段缺失 / 值为空的文档一律排到最后,不干扰前排结论
### 常用命令
`search-yaml.py` 用于扫**带 frontmatter 的产物**——feature spec / issue spec / requirements / adrs / guides / library-docs。
`.codestable/compound/` 由 `cs-keep` 写纯 markdown(无 frontmatter),**不用 search-yaml**,用全文检索即可——grep / ripgrep / 框架自带搜索都行,下例以 grep 示意:
```bash
grep -r "关键词" .codestable/compound/
grep -rl "prisma" .codestable/compound/ # 只列文件名
ls -lt .codestable/compound/ | head # 看最近沉淀
```
带 frontmatter 的目录用 search-yaml:
```bash
# 搜索 feature 方案 doc
python .codestable/tools/search-yaml.py --dir .codestable/features --filter doc_type=feature-design --filter status=approved
# 按时间排序
python .codestable/tools/search-yaml.py --dir .codestable/library-docs --sort-by last_reviewed --order asc # 最久没 review 的在前(找陈旧文档)
python .codestable/tools/search-yaml.py --dir .codestable/guides --filter status=current --sort-by last_reviewed --order asc
# 输出控制
python .codestable/tools/search-yaml.py --dir .codestable/features --filter status=approved --full
python .codestable/tools/search-yaml.py --dir .codestable/features --filter tags~=llm --json
```
### 典型使用场景
| 场景 | 命令建议 |
|---|---|
| feature-design 开始前查 compound 已有沉淀 | `grep -r "{关键词}" .codestable/compound/` |
| issue-analyze 根因分析前查历史 | `grep -rl "{关键词}" .codestable/compound/` 再人工挑相关的看 |
| cs-keep 落盘前查重叠 | `grep -rl "{关键词}" .codestable/compound/`,命中就先看那条决定更新还是新写 |
| 找最久没 review 的库文档 / 指南 | `--dir {目录} --filter status=current --sort-by last_reviewed --order asc` |
---
## 2. validate-yaml.py
YAML 语法校验工具。用于验证 frontmatter 语法和必填字段。
```bash
# 校验单个文件的 YAML 语法
python .codestable/tools/validate-yaml.py --file {文件路径} --yaml-only
# 校验必填字段
python .codestable/tools/validate-yaml.py --file {文件路径} --require doc_type --require status
# 批量校验目录下所有文件
python .codestable/tools/validate-yaml.py --dir {目录} --require doc_type --require status
```
-338
View File
@@ -1,338 +0,0 @@
#!/usr/bin/env python3
"""
search-yaml.py — Generic YAML-frontmatter search tool for markdown document directories.
Works on any directory of .md files that use YAML frontmatter (--- ... ---).
Designed for AI agent use: fast, structured output, no required external dependencies.
Filter syntax (--filter flag, repeatable, AND logic):
key=value Exact match on a scalar field (case-insensitive)
key=a|b Exact match against any candidate value (OR)
key~=value Substring match on a string field, or element-in for list fields
key~=a|b Substring/list match against any candidate value (OR)
Usage examples:
# Search feature specs by status
python .codestable/tools/search-yaml.py --dir .codestable/features --filter doc_type=feature-design --filter status=approved
# Filter by tag (list element match) and full-text search body + frontmatter values
python .codestable/tools/search-yaml.py --dir .codestable/features --filter tags~=prisma
python .codestable/tools/search-yaml.py --dir .codestable/features --query "shadow database"
# JSON output for AI agent consumption
python .codestable/tools/search-yaml.py --dir .codestable/issues --filter status=open --json
# Sort by a frontmatter date field (works on any ISO-8601 date string, YAML date, or sortable value)
python .codestable/tools/search-yaml.py --dir .codestable/library-docs --sort-by last_reviewed --order asc # oldest first (stalest)
# NOTE: .codestable/compound/ is plain markdown (no frontmatter) — use grep instead:
# grep -r "keyword" .codestable/compound/
# Works on any yaml-frontmatter markdown directory
python .codestable/tools/search-yaml.py --dir docs/decisions --filter status=accepted
python .codestable/tools/search-yaml.py --dir content/posts --filter tags~=python --query "asyncio"
"""
import argparse
import json
import sys
from pathlib import Path
try:
import yaml # type: ignore
_HAS_PYYAML = True
except ImportError:
_HAS_PYYAML = False
# ---------------------------------------------------------------------------
# Frontmatter parsing (PyYAML used when available, builtin fallback otherwise)
# ---------------------------------------------------------------------------
def _parse_yaml_scalar(val: str):
val = val.strip()
if val.startswith("[") and val.endswith("]"):
inner = val[1:-1]
return [item.strip().strip("'\"") for item in inner.split(",") if item.strip()]
lower = val.lower()
if lower in ("true", "yes"):
return True
if lower in ("false", "no"):
return False
if lower in ("null", "~", ""):
return None
return val
def parse_frontmatter(text: str) -> tuple[dict, str]:
"""
Split a markdown document into (frontmatter_dict, body_text).
Returns ({}, full_text) when no frontmatter is present.
"""
if not text.startswith("---"):
return {}, text
end = text.find("\n---", 3)
if end == -1:
return {}, text
fm_text = text[3:end].strip()
body = text[end + 4:].strip()
if _HAS_PYYAML:
try:
meta = yaml.safe_load(fm_text)
return (meta or {}), body
except yaml.YAMLError:
# Malformed frontmatter — fall through to the lenient builtin parser
# so partial / hand-written frontmatter still produces best-effort results.
pass
# Minimal fallback: handles scalar values and inline lists
meta: dict = {}
for line in fm_text.splitlines():
if not line.strip() or line.startswith("#") or ":" not in line:
continue
key, _, raw = line.partition(":")
meta[key.strip()] = _parse_yaml_scalar(raw)
return meta, body
# ---------------------------------------------------------------------------
# Document loading
# ---------------------------------------------------------------------------
def load_documents(directory: Path) -> list[dict]:
docs = []
for md_file in sorted(directory.rglob("*.md")):
try:
text = md_file.read_text(encoding="utf-8")
except OSError as exc:
print(f"[warn] Cannot read {md_file.name}: {exc}", file=sys.stderr)
continue
meta, body = parse_frontmatter(text)
docs.append({
"file": str(md_file.relative_to(directory)),
"path": str(md_file),
"meta": meta,
"body": body,
})
return docs
# ---------------------------------------------------------------------------
# Filter parsing and evaluation
# ---------------------------------------------------------------------------
def _split_filter_values(value: str) -> list[str]:
values = [part.strip() for part in value.split("|")]
return [part for part in values if part] or [value.strip()]
class Filter:
"""Parsed representation of a single --filter expression."""
def __init__(self, raw: str):
if "~=" in raw:
key, _, value = raw.partition("~=")
self.key = key.strip()
self.value = value.strip()
self.values = _split_filter_values(self.value)
self.operator = "contains"
elif "=" in raw:
key, _, value = raw.partition("=")
self.key = key.strip()
self.value = value.strip()
self.values = _split_filter_values(self.value)
self.operator = "exact"
else:
raise argparse.ArgumentTypeError(
f"Invalid filter expression {raw!r}. "
"Use 'key=value' for exact match or 'key~=value' for substring/list-contains match. "
"Use pipes for OR values, e.g. 'status=approved|draft'."
)
def matches(self, meta: dict) -> bool:
field_val = meta.get(self.key)
if field_val is None:
return False
if self.operator == "exact":
return any(str(field_val).lower() == value.lower() for value in self.values)
# contains: substring for strings, element-in for lists
if isinstance(field_val, list):
return any(
value.lower() == str(item).lower()
for value in self.values
for item in field_val
)
return any(value.lower() in str(field_val).lower() for value in self.values)
def __repr__(self):
op = "~=" if self.operator == "contains" else "="
return f"Filter({self.key}{op}{self.value})"
def parse_filter(raw: str) -> Filter:
"""argparse type converter for --filter."""
return Filter(raw)
_MISSING = object()
def _sort_key(doc: dict, field: str):
"""
Sort key for --sort-by. Docs missing the field sort to the end regardless
of --order. Dates (datetime.date / datetime.datetime) and strings are both
normalized to their string form — ISO 8601 date strings sort the same
lexicographically as YAML-parsed date objects' isoformat().
"""
val = doc["meta"].get(field, _MISSING)
if val is _MISSING or val is None:
return (1, "")
try:
return (0, val.isoformat()) # datetime.date / datetime.datetime
except AttributeError:
return (0, str(val))
def doc_matches(doc: dict, filters: list[Filter], query: str | None) -> bool:
meta = doc["meta"]
for f in filters:
if not f.matches(meta):
return False
if query:
needle = query.lower()
haystack = doc["body"].lower() + " " + " ".join(str(v) for v in meta.values()).lower()
if needle not in haystack:
return False
return True
# ---------------------------------------------------------------------------
# Output formatting
# ---------------------------------------------------------------------------
def _meta_summary(meta: dict) -> str:
"""One-line summary of frontmatter fields, skipping slug/date for brevity."""
skip = {"slug"}
parts = []
for k, v in meta.items():
if k in skip:
continue
if isinstance(v, list):
parts.append(f"{k}=[{', '.join(str(i) for i in v)}]")
else:
parts.append(f"{k}={v}")
return " ".join(parts)
def format_summary(doc: dict) -> str:
return f"### {doc['file']}\n{_meta_summary(doc['meta'])}"
def format_full(doc: dict) -> str:
return format_summary(doc) + "\n\n" + doc["body"]
def print_text(results: list[dict], full: bool) -> None:
print(f"Found {len(results)} document(s).\n")
sep = "\n" + "─" * 60 + "\n"
chunks = [format_full(d) if full else format_summary(d) for d in results]
print(sep.join(chunks))
def print_json(results: list[dict], full: bool) -> None:
output = []
for doc in results:
body = doc["body"]
if not full and len(body) > 400:
body = body[:400] + "…"
output.append({"file": doc["file"], "meta": doc["meta"], "body": body})
print(json.dumps(output, ensure_ascii=False, indent=2))
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def _build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Generic YAML-frontmatter search across a directory of markdown files.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("--dir", metavar="DIR", required=True,
help="Directory of .md files to search.")
parser.add_argument("--filter", "-f", metavar="EXPR", dest="filters",
type=parse_filter, action="append", default=[],
help="Frontmatter filter expression. Repeatable (AND logic). "
"key=value for exact match; key~=value for substring (strings) or element-in (lists). "
"Use pipes for OR values, e.g. key=a|b.")
parser.add_argument("--query", "-q", metavar="TEXT",
help="Full-text search in document body and frontmatter values.")
parser.add_argument("--full", action="store_true",
help="Print full document body instead of just the frontmatter summary.")
parser.add_argument("--json", dest="as_json", action="store_true",
help="Output results as a JSON array.")
parser.add_argument("--sort-by", metavar="FIELD", dest="sort_by",
help="Sort results by a frontmatter field (e.g. last_reviewed, date, updated_at). "
"ISO-8601 date strings and YAML-parsed dates both sort correctly. "
"Docs missing the field are pushed to the end.")
parser.add_argument("--order", choices=("asc", "desc"), default="desc",
help="Sort order when --sort-by is set. Default: desc (newest first).")
return parser
def _resolve_directory(dir_arg: str) -> Path:
directory = Path(dir_arg)
if not directory.exists():
print(f"[error] Directory not found: {directory}", file=sys.stderr)
sys.exit(1)
if not directory.is_dir():
print(f"[error] Not a directory: {directory}", file=sys.stderr)
sys.exit(1)
return directory
def _sort_results(results: list[dict], sort_by: str, order: str) -> list[dict]:
def has_field(d: dict) -> bool:
return sort_by in d["meta"] and d["meta"][sort_by] is not None
present = [d for d in results if has_field(d)]
missing = [d for d in results if not has_field(d)]
present.sort(key=lambda d: _sort_key(d, sort_by), reverse=(order == "desc"))
return present + missing
def main() -> None:
args = _build_parser().parse_args()
directory = _resolve_directory(args.dir)
docs = load_documents(directory)
if not docs:
print(f"No .md files found in {directory}")
return
results = [d for d in docs if doc_matches(d, args.filters, args.query)]
if not results:
print("No matching documents found.")
return
if args.sort_by:
results = _sort_results(results, args.sort_by, args.order)
if args.as_json:
print_json(results, full=args.full)
else:
print_text(results, full=args.full)
if __name__ == "__main__":
main()
-315
View File
@@ -1,315 +0,0 @@
#!/usr/bin/env python3
"""
validate-yaml.py — Validate YAML frontmatter syntax in markdown files.
Scans markdown files for YAML frontmatter (--- ... ---) and checks:
1. Frontmatter block is properly delimited (opening and closing ---)
2. YAML syntax is valid (parseable without errors)
3. (Optional) Required fields are present (--require flag)
Designed for AI agent use: structured output, exit code reflects pass/fail,
no required external dependencies (falls back to builtin parser if PyYAML unavailable).
Usage examples:
# Validate all .md files under codestable/features
python codestable/tools/validate-yaml.py --dir codestable/features
# Validate a single file
python codestable/tools/validate-yaml.py --file codestable/features/2026-04-11-auth/auth-design.md
# Check that required fields exist in frontmatter
python codestable/tools/validate-yaml.py --dir codestable/features --require doc_type --require status
# JSON output for programmatic consumption
python codestable/tools/validate-yaml.py --dir docs/api --json
# Validate the doc-api manifest
python codestable/tools/validate-yaml.py --file docs/api/manifest.yaml --yaml-only
"""
import argparse
import json
import sys
from pathlib import Path
# Force UTF-8 stdout/stderr on Windows where default codepage (e.g. GBK / cp936)
# can't encode the ✓ / ✗ icons used in text output. Safe no-op on POSIX.
# Streams that aren't a real TextIOWrapper (e.g. captured by pytest, redirected
# through some IDEs) raise io.UnsupportedOperation — a ValueError + OSError
# subclass — and we just leave the original encoding in place.
for _stream in (sys.stdout, sys.stderr):
if hasattr(_stream, "reconfigure"):
try:
_stream.reconfigure(encoding="utf-8")
except (OSError, ValueError):
pass
# ---------------------------------------------------------------------------
# YAML parsing
# ---------------------------------------------------------------------------
_HAS_PYYAML = False
try:
import yaml # type: ignore
_HAS_PYYAML = True
except ImportError:
pass
def _builtin_parse_yaml(text: str) -> dict:
"""Minimal YAML parser for flat key-value frontmatter (no nested structures)."""
result: dict = {}
for line in text.splitlines():
stripped = line.strip()
if not stripped or stripped.startswith("#") or ":" not in stripped:
continue
key, _, raw = stripped.partition(":")
val = raw.strip()
# Inline list
if val.startswith("[") and val.endswith("]"):
inner = val[1:-1]
result[key.strip()] = [
item.strip().strip("'\"") for item in inner.split(",") if item.strip()
]
else:
result[key.strip()] = val.strip("'\"") if val else ""
return result
def parse_yaml_text(text: str) -> tuple[dict | None, str | None]:
"""
Parse a YAML string. Returns (parsed_dict, None) on success,
or (None, error_message) on failure.
"""
if _HAS_PYYAML:
try:
result = yaml.safe_load(text)
if result is None:
return {}, None
if not isinstance(result, dict):
return None, f"Expected a mapping, got {type(result).__name__}"
return result, None
except yaml.YAMLError as exc:
return None, str(exc)
else:
# Builtin fallback — can only detect gross syntax issues
try:
result = _builtin_parse_yaml(text)
return result, None
except Exception as exc:
return None, str(exc)
# ---------------------------------------------------------------------------
# Frontmatter extraction
# ---------------------------------------------------------------------------
def extract_frontmatter(text: str) -> tuple[str | None, str | None]:
"""
Extract YAML frontmatter from a markdown file.
Returns (frontmatter_text, None) on success,
or (None, error_message) if frontmatter is missing or malformed.
"""
if not text.startswith("---"):
return None, "No opening '---' delimiter found"
end = text.find("\n---", 3)
if end == -1:
return None, "No closing '---' delimiter found (frontmatter block not terminated)"
fm_text = text[3:end].strip()
if not fm_text:
return None, "Frontmatter block is empty"
return fm_text, None
# ---------------------------------------------------------------------------
# Validation logic
# ---------------------------------------------------------------------------
class ValidationResult:
def __init__(self, file_path: str):
self.file = file_path
self.errors: list[str] = []
self.warnings: list[str] = []
self.fields: list[str] = [] # fields found in frontmatter
@property
def ok(self) -> bool:
return len(self.errors) == 0
def to_dict(self) -> dict:
d: dict = {"file": self.file, "status": "pass" if self.ok else "fail"}
if self.errors:
d["errors"] = self.errors
if self.warnings:
d["warnings"] = self.warnings
if self.fields:
d["fields"] = self.fields
return d
def _check_required(parsed: dict | None, required_fields: list[str] | None, result: ValidationResult) -> None:
if not required_fields:
return
for field in required_fields:
if field not in (parsed or {}):
result.errors.append(f"Missing required field: '{field}'")
def _warn_if_builtin(result: ValidationResult) -> None:
if not _HAS_PYYAML:
result.warnings.append(
"PyYAML not installed — using builtin fallback parser "
"(may miss some syntax errors). Install with: pip install pyyaml"
)
def _validate_file(
file_path: Path,
required_fields: list[str] | None,
base_dir: Path | None,
mode: str, # "markdown" | "yaml"
) -> ValidationResult:
display_path = str(file_path.relative_to(base_dir)) if base_dir else str(file_path)
result = ValidationResult(display_path)
try:
text = file_path.read_text(encoding="utf-8")
except OSError as exc:
result.errors.append(f"Cannot read file: {exc}")
return result
if mode == "markdown":
yaml_text, extract_err = extract_frontmatter(text)
if extract_err:
result.errors.append(extract_err)
return result
else:
yaml_text = text
parsed, parse_err = parse_yaml_text(yaml_text)
if parse_err:
result.errors.append(f"YAML syntax error: {parse_err}")
return result
result.fields = list(parsed.keys()) if parsed else []
_check_required(parsed, required_fields, result)
_warn_if_builtin(result)
return result
def validate_markdown_file(file_path, required_fields=None, base_dir=None):
"""Validate YAML frontmatter in a single markdown file."""
return _validate_file(file_path, required_fields, base_dir, "markdown")
def validate_yaml_file(file_path, required_fields=None, base_dir=None):
"""Validate a pure YAML file (not markdown with frontmatter)."""
return _validate_file(file_path, required_fields, base_dir, "yaml")
# ---------------------------------------------------------------------------
# Output
# ---------------------------------------------------------------------------
def print_text_results(results: list[ValidationResult]) -> None:
passed = sum(1 for r in results if r.ok)
failed = len(results) - passed
print(f"Validated {len(results)} file(s): {passed} passed, {failed} failed.\n")
for r in results:
icon = "✓" if r.ok else "✗"
print(f" {icon} {r.file}")
for err in r.errors:
print(f" ERROR: {err}")
for warn in r.warnings:
print(f" WARN: {warn}")
if failed > 0:
print(f"\n{failed} file(s) have YAML errors.")
else:
print("\nAll files valid.")
def print_json_results(results: list[ValidationResult]) -> None:
output = {
"total": len(results),
"passed": sum(1 for r in results if r.ok),
"failed": sum(1 for r in results if not r.ok),
"results": [r.to_dict() for r in results],
}
print(json.dumps(output, indent=2, ensure_ascii=False))
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def _build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Validate YAML frontmatter in markdown files or pure YAML files.",
formatter_class=argparse.RawDescriptionHelpFormatter,
)
source = parser.add_mutually_exclusive_group(required=True)
source.add_argument("--dir", type=str, help="Directory to scan recursively for .md files")
source.add_argument("--file", type=str, help="Single file to validate")
parser.add_argument("--require", action="append", default=[], metavar="FIELD",
help="Require this field in frontmatter (repeatable)")
parser.add_argument("--json", action="store_true", dest="json_output",
help="Output results as JSON")
parser.add_argument("--yaml-only", action="store_true",
help="Treat input as pure YAML (not markdown with frontmatter). "
"Use for .yaml/.yml files like manifest.yaml.")
return parser
def _validate_single(path_str: str, require: list[str], yaml_only: bool) -> list[ValidationResult]:
fp = Path(path_str)
if not fp.exists():
print(f"Error: File not found: {fp}", file=sys.stderr)
sys.exit(2)
if yaml_only or fp.suffix in (".yaml", ".yml"):
return [validate_yaml_file(fp, require)]
return [validate_markdown_file(fp, require)]
def _validate_directory(dir_str: str, require: list[str]) -> list[ValidationResult]:
dp = Path(dir_str)
if not dp.is_dir():
print(f"Error: Directory not found: {dp}", file=sys.stderr)
sys.exit(2)
md_files = sorted(dp.rglob("*.md"))
yaml_files = sorted(dp.rglob("*.yaml")) + sorted(dp.rglob("*.yml"))
if not md_files and not yaml_files:
print(f"No .md or .yaml files found under {dp}", file=sys.stderr)
sys.exit(2)
results = [validate_markdown_file(md, require, dp) for md in md_files]
results += [validate_yaml_file(yf, require, dp) for yf in yaml_files]
return results
def main() -> None:
args = _build_parser().parse_args()
if args.file:
results = _validate_single(args.file, args.require, args.yaml_only)
else:
results = _validate_directory(args.dir, args.require)
if args.json_output:
print_json_results(results)
else:
print_text_results(results)
sys.exit(0 if all(r.ok for r in results) else 1)
if __name__ == "__main__":
main()