* feat(model-files): sync precision and quant type lists Bring the code-side fallback tables in line with the precisions and quant types actually in use. `int8` and `mxfp8` are already live in the DB row but were missing from fpQualityRank, so they ranked 0 and sorted below `nf4` in download ordering. Adds Unsloth dynamic (`*_K_XL`), extra imatrix (`IQ4_KS`, `IQ3_M`, `IQ3_S`) and ternary (`TQ2_0`, `TQ1_0`) GGUF quants, plus a `None` quant type meaning "not quantized" so an unquantized GGUF has a valid selection. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(model-files): rank Q3_K_XL above Q3_K_L Unsloth dynamic _XL quants are larger and higher quality than the _L variant of the same bit width, so within a bit width the order is XL > L > M > S. Q3_K_XL was sitting below Q3_K_L. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(model-files): tighten None quant-type handling Review follow-ups on the None sentinel: - FileInfo showed "Quant Type: None" raw; now renders "Unquantized" via a shared formatQuantType helper. - The account "Preferred Quant Type" select offered None, which reads as "no preference" but stores "prefer unquantized" and feeds getPrimaryFile. Excluded — leaving the preference unset is how you express no preference. - Switching quant off None left metadata.fp set but hidden and un-editable, where it still scored in getPrimaryFile and keyed checkConflictingFiles. Both forms now clear fp when the new quant isn't None. - Nothing required a precision when quant was None, so an unquantized GGUF could save with less information than before the None option existed. Also documents the sentinel's semantics — notably that DELETEing None from the DB list makes the Precision field unreachable for GGUF uploads with no error surfaced — and that array order is UI order, so a reordering rollout needs PUT rather than POST. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(model-files): resolve merge fp by key presence, not ?? The merge mapper resolved effectiveFp with `??`, which treats the explicit `fp: null` written when quant moves off None as absent and falls back to the source file's original precision. The Precision field then displayed a value the merge would not save, and could produce a quantType=None file with no fp — the state the upload refinement exists to prevent, on a path that refinement doesn't cover. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
5.8 KiB
Spec: DB-managed model file precisions & quant types
ClickUp: 868k69pey
Goal
Move the two hardcoded model-file metadata lists out of constants.ts and into the DB so mods
can edit them without a deploy (these lists grow over time as new GGUF quant types / precisions appear).
- precisions (
modelFileFp): 11 values, best-to-worstfp32 … int4 - quantTypes (
modelFileQuantTypes): 34 values —None(unquantized) plus theQ*/IQ*/TQ*set
Storage
Single KeyValue row, key modelFileOptions:
{ "precisions": ["fp32", "fp16", ...], "quantTypes": ["None", "Q8_0", ...] }
The hardcoded constants.modelFileFp / constants.modelFileQuantTypes stay as the default/fallback
when the key is absent or unreadable (same layered pattern as update-user-score multipliers:
DB → hardcoded default).
Array order is UI order. The dropdowns render the stored arrays verbatim, so both lists are kept
best-quality-first. This is also why a rollout that reorders values must use PUT (wholesale replace)
rather than POST (which preserves existing order and appends new values to the tail).
The None quant type
None is a sentinel meaning "not quantized", not a quantization. It exists so an unquantized GGUF has
a valid answer to the .gguf-requires-a-quantType rule without duplicating the precision field as
F16/F32 quant values. It carries behavior that plain list values do not:
- Ranks above every quantized value in
quantQualityRank(unquantized is the highest quality). - Selecting it is the only way to reach the Precision field on a GGUF upload — the file editor and
the merge-versions mapper show Precision instead of hiding it when quant is
None, and a refinement then requires a precision. DeletingNonefrom the DB list silently makes Precision unreachable for every GGUF, with no error surfaced anywhere. Do notDELETEit. - Never rendered literally: display helpers show "Unquantized" (
formatQuantType), and it is filtered out of the account-level "Preferred Quant Type" select, where "no preference" is expressed by leaving the preference unset.
Service
Folded into src/server/services/model-file.service.ts (alongside the other model-file CRUD).
Uses dbKV (~/server/db/db-helpers) for KeyValue access — the repo convention (cf. system.router
getDbKV, training.router). No app-layer cache: caching lives at the edge (edgeCacheIt on the
public procedure); dbKV is bound to the primary, so a write is self-consistent on the next read.
getModelFileOptions()→dbKV.get(KEY), normalized; falls back to constants defaults when absent.setModelFileOptions/addModelFileOptions/removeModelFileOptions({ precisions?, quantTypes? }) → read-before-write via a sharedmutateModelFileOptions(input, merge)helper, thendbKV.set.
Mod management endpoint (webhook, token-secured)
src/pages/api/admin/model-file-options.ts using WebhookEndpoint (?token=$WEBHOOK_TOKEN).
Method-based REST (no action param) — the verb says what it does; body { precisions?, quantTypes? }:
GET→ current{ precisions, quantTypes }(live, bypasses cache)PUT→ replace the provided list(s) wholesalePOST→ add value(s) to the existing list(s)DELETE→ remove value(s) from the existing list(s)
Writes read-before-write from the primary (dbWrite) so a partial mutation can't clobber the
preserved list during replication lag. Body validated: non-empty string[], ≥1 of the two keys.
Public read (edge-cached, client-facing)
tRPC modelFile.getOptions (src/server/routers/model-file.router.ts), publicProcedure with
.use(edgeCacheIt({ ttl: CacheTTL.sm, tags: () => [MODEL_FILE_OPTIONS_EDGE_TAG] })) → 3-min CDN
s-maxage, tagged for purge. Returns { precisions, quantTypes } from the service. Established
edge-cache convention (cf. system.router getDbKV, generation.router tag+purge); the repo's
tRPC client uses non-batched httpLink GET so edgeCacheIt applies.
Cache busting: every mod write (mutateModelFileOptions) calls
purgeCache({ tags: [MODEL_FILE_OPTIONS_EDGE_TAG] }) after dbKV.set, so the CDN serves fresh
immediately. Already-loaded browser tabs still refetch on their 3-min React Query staleTime.
purgeCache no-ops without CF_ZONE_ID (dev).
Client — dropdowns
Shared hook src/hooks/useModelFileOptions.ts: trpc.modelFile.getOptions.useQuery (staleTime ~3min
to match edge cache), returning { precisions, quantTypes }, falling back to constants.* while
loading / on error so dropdowns never render empty.
Wired into all three consumers:
src/components/Resource/Files.tsx— the model file editor (Quant + Precision selects)src/components/Model/Actions/MergeVersions.tsx— the file mapper's Quant + Precision selectssrc/components/Account/SettingsCard.tsx— "Preferred Precision" / "Preferred Quant Type"
Server validation (the one real design decision)
The 4 zod schemas use z.enum(constants.modelFileFp) / z.enum(constants.modelFileQuantTypes). z.enum
is static at module-load, so a newly-added DB value would FAIL upload/download validation — defeating the
feature. Options:
- A (recommended): relax those fields to
z.string()(nullish, with a sane.max()); values are mod-curated + selected from the editor dropdown, low-risk metadata tags. - B: keep
z.enumas a superset — new values require also editing constants (defeats the purpose). - C: async-refine each schema against the cached DB list (correct but heavy; schemas are imported in sync contexts).
Out of scope (MVP)
- No in-app mod UI; management is via the webhook endpoint only (matches "like the KoN queues").
- The 4 zod schemas relax to
z.string()(decision A) — values are mod-curated + dropdown-selected.