8197 Commits

Author SHA1 Message Date
Ziyang Guo
75363af54b fix(mistral): honor dataset language in figure prompts (#18021)
### Summary

Refs #17885.

Mistral figure enrichment now receives the dataset language through the
production parsing path. `by_mistral_ocr` forwards `lang` to
`MistralParser.parse_pdf`; the parser stores the normalized language and
passes it to the figure-description prompt. Empty or missing values
still fall back to English.
nightly
2026-08-11 20:38:24 +08:00
Carl Calaquian
ad6fdfd7b4 fix(sandbox): update Tenki SDKs and drop removed project_id (#18117)
### Summary

Brings both halves of the Tenki sandbox provider onto current SDKs and
removes `project_id`, which Tenki deleted from its API.

**Go:** `github.com/LuxorLabs/tenki-sdk-go/sandbox` `v0.5.2` → `v0.7.0`
(current latest).
**Python:** the provider's SDK was renamed on PyPI — `tenki-sandbox` is
frozen at 0.4.0 and everything from 0.5 ships as
[`tenki`](https://pypi.org/project/tenki/). The docs told operators to
`pip install tenki-sandbox`, which installs a stale SDK that no longer
matches this provider's expectations.

**`project_id` is gone.** Tenki removed project scoping from the sandbox
API in 0.5.x: `Client.create()` no longer accepts `project_id`, so the
current code path would raise `TypeError` against a current SDK. It was
also marked `required: True` in the config schema, so the Admin >
Sandbox Settings form asked for a value that no longer exists.
2026-08-11 20:37:21 +08:00
balibabu
ec63ab37b1 Fix: Remove the model from the wiki template. (#18119) 2026-08-11 20:29:08 +08:00
balibabu
1d81ca27cc Fix: A dataset-level tree can also view file-level data. (#18123) 2026-08-11 20:28:44 +08:00
buua436
0cfd8f41e4 fix: improve incremental wiki compilation (#18130)
### What problem does this PR solve?

Incremental Wiki compilation could lose provenance for claim-light
entities, produce unstable page groups across embedding models, route
entities to unrelated pages, and assign topics without sufficient
page-level context. Document removals and page membership changes could
also leave stale Wiki state.

This PR:

- preserves source document and chunk provenance throughout entity
matching, reduction, page generation, and deletion;
- uses embeddings to retrieve candidates and the LLM to make final page
grouping and incremental routing decisions;
- batches embedding and LLM operations with bounded concurrency and
deterministic fallbacks;
- selects source-scoped topic candidates with embeddings before the page
LLM chooses the final topic;
- rebuilds Wiki state when the compilation mode or embedding model
changes;
- normalizes Wiki array fields returned by the API and retains entities
without relations in graph responses.

### Type of change

- [x] Bug Fix (non-breaking change which fixes an issue)
2026-08-11 20:13:04 +08:00
Jin Hai
4386ff71b0 Go: fix context, part2 (#18133)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 20:12:29 +08:00
Jack
93dca789b5 refactor(parser): make markdown golden meta-driven, drop generator script (#18125) 2026-08-11 19:53:56 +08:00
Haruko386
3ef6e45e43 fix: return error when chat-channel start has error (#18068) 2026-08-11 19:36:35 +08:00
Jin Hai
c75edbfbe8 Go: fix context (#18118)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 19:19:29 +08:00
nikminer
8bd5768ebc Integrate MWS model with API support and enhance chat functionality (#17959)
## What

This pull request adds **MWS GPT Model Hub** as a built-in model
provider in RAGFlow.

The integration allows users to configure an MWS project endpoint and
token, discover the models available to that project, and use supported
MWS models for chat completion, embeddings, and reranking.

Co-authored-by: ilarionov_n <ilarionov_n@promis.ru>
2026-08-11 19:12:42 +08:00
chanx
5e2c0eee28 refactor: use CopyToClipboard in embed-app-modal (#18115) 2026-08-11 19:08:33 +08:00
chanx
dbe2bf8b8b fix: sanitize img tags in markdown via shared SafeImg component (#18112) 2026-08-11 19:08:15 +08:00
Lynn
cd6996b301 Fix: xinference asr (#18110) 2026-08-11 19:07:50 +08:00
Yingfeng
fbcb8656ca Revert "Refine agentic search & orchestration loop" (#18108) 2026-08-11 18:53:49 +08:00
Jack
4a8bb1b72a test(parser): add shared golden-doc + alignment helpers (#18098)
Centralize the shared golden-doc + alignment helpers in `align_test.go` so the format-specific PRs (text&code, markdown golden, HTML) reuse one implementation instead of each carrying their own copy of the scaffolding.
2026-08-11 18:00:25 +08:00
dependabot[bot]
f0d1ff52ab chore(deps): bump fast-uri from 3.1.0 to 3.1.5 in /web (#18103)
Bumps [fast-uri](https://github.com/fastify/fast-uri) from 3.1.0 to

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-11 17:52:30 +08:00
qinling0210
d49af7f218 Generate navigation, navigation search (#18096)
### Summary
2 API

POST /api/v1/datasets/{dataset_id}/navigation

GET
/api/v1/datasets/{dataset_id}/navigation/search?q={query}&mode={mode}&top_k={topk}


2 cli

uv run --no-sync python3 admin/client/ragflow_cli.py -h 127.0.0.1 -p
9380 -t user

ragflow> GENERATE NAVIGATION OF DATASET 'frames tree';

ragflow> NAVIGATION SEARCH 'Christie introduced blockchain-based digital
passports' IN DATASET 'frames tree' MODE 'all' topk 20;

(mode: chunk, nav_cluster, nav_doc, navigation_tree, all)
2026-08-11 17:48:24 +08:00
Jack
a399b93143 Test(chunker): add golden parity harness, fixtures, and live Go<->Python tool (#17735)
Golden-parity test infrastructure for the **Go `TokenChunker` ↔ Python alignment**.
It runs the Go chunker over a committed case set (`testdata/parity/cases/`) and diffs each output against a captured Python golden (`testdata/parity/golden/`), honoring a `known_diffs.json` ratchet (`extra_fields` / `chunk_count` / `chunk_text`) so accepted divergences are tracked rather than silently widening.
2026-08-11 17:46:34 +08:00
chanx
ba671861e2 Fix: stabilize agent log handleSearch and preserve page_size on reset (#18087) 2026-08-11 17:25:24 +08:00
balibabu
7793d3c01c Fix: Retrieve the latest timeline data after cancelling the timeline task. (#18102) 2026-08-11 17:25:09 +08:00
balibabu
5dde775e19 Fix: If a message contains multiple files, they need to be displayed with automatic line wrapping. (#18109) 2026-08-11 17:24:47 +08:00
Lynn
59a524de70 Fix: remove some deprecated model (#18105) 2026-08-11 17:24:23 +08:00
Wang Qi
8e04f95773 Fix to force logout user when set user as inactive (#893) (#18104) 2026-08-11 17:00:57 +08:00
Jin Hai
d7661b676d Go: fix context and warnings (#18097)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 16:18:49 +08:00
chanx
2409248801 fix: show filtered session count in session list header (#18094) 2026-08-11 16:14:08 +08:00
chanx
d33f87d3ca feat: rename agent log export button to export current page (#18090) 2026-08-11 16:13:56 +08:00
euvre
82d038543a fix: wrap long log messages in agent pipeline log sheet (#17782) 2026-08-11 16:08:43 +08:00
Haruko386
dd52dee600 feat[Go]: monitoring NATs and refactoring concurrency logic (#18049)
### Summary

As title
2026-08-11 14:36:11 +08:00
Jin Hai
25b579ac6c TS: more license declarations (#18081)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 14:21:14 +08:00
balibabu
61fa96ee06 Fix: Click the blank area to prevent UpdateLogSheet from closing. (#18085) 2026-08-11 14:09:54 +08:00
balibabu
398d81f587 Fix: Move the content of the "Token Size" tab from the token chunker operator to the "Delimiter" tab. #17888 (#17975) 2026-08-11 14:07:15 +08:00
euvre
d33b937be6 fix(web): prevent column shift in data source log table (#17940) 2026-08-11 13:59:51 +08:00
buua436
5fd66154f3 fix: improve structure keyword search (#18055) 2026-08-11 13:58:00 +08:00
Zhichang Yu
64533e5b5e Refactor splitByTokens and wire wiki incremental compile (#18083)
Port dataset-level wiki incremental compile and refactor splitByTokens
token budgeting. Includes replace-only wiki merge, KNN dedup routing,
and template/config wiring.
2026-08-11 13:46:55 +08:00
balibabu
f83c4f97c1 Fix: After creating a wiki template and deleting the model instance, the default model appears empty, but it can still be edited and saved successfully. (#18075) 2026-08-11 13:45:04 +08:00
Yingfeng
f1641228e2 Refine agentic search & orchestration loop (#18057)
## Summary

This PR improves the RAGFlow agentic-search path in three areas: it
stops the outer agent from re-looping over the same rag call, lets the
medium thinking mode discover and follow new sub-claims mid-loop, and
strengthens retrieval by having the LLM emit synonym-rich queries with
time/date/number terms boosted.

1. Avoid the outer re-loop — keep all multi-hop cycles inside agentic
RAG

2. Dynamic claims in medium mode — keep querying newly discovered
sub-questions
medium now enables allows_dynamic_claims. During orchestration, when
claim analysis discovers a new required sub-question
(discovered_claims), the loop spawns it as a new ClaimTarget and
continues searching it in subsequent cycles (bounded by the
dynamic-claim budget) instead of stopping. Also added:

3. Stronger query strategy — synonym-rich queries + time/date/number
weighting

LLM-generated synonyms: the claim-analysis prompt now instructs the
model to write each next_queries entry as a retrieval-boosted query that
actively folds in entity aliases, DATE/TIME synonyms (e.g. 1994 → 1994,
66th Academy Awards), and number/unit variants (e.g. 1.95 m → 6 ft 5
in).

Time/date/number boosting: query.py boosts numeric/date tokens to a high
weight (_NUM_DATE_TOKEN_RE).
2026-08-11 13:40:11 +08:00
chanx
f930b1c1bc Fix: New sessions will automatically carry over messages from the previous session by default. (#18078) 2026-08-11 13:38:46 +08:00
chanx
e0d199ea7b fix(floating-selection-toolbar): adjust dropdown direction based on available space (#18053) 2026-08-11 13:31:42 +08:00
chanx
11c9649a6c Fixed: Updated the text color style of the copy button. (#18062) 2026-08-11 13:30:57 +08:00
Wang Qi
bd876c796e Fix: restrict xxx_ids to limit 100 (#18074) 2026-08-11 13:28:52 +08:00
Wang Qi
d7d6bb5b6a Fix max chunk_ids = 100 (#18080) 2026-08-11 13:13:08 +08:00
Jin Hai
fc44d2fe3f Go: fix context (#18076)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
dev-20260811-2
2026-08-11 11:54:57 +08:00
jay77721
9f0663d4d0 fix(ingestion): avoid duplicate chunk text injection in Extractor prompts (#18034) 2026-08-11 11:53:34 +08:00
chanx
d199f4b867 fix(search): pass selected pageSize when retrieving chunks on document change (#18072) 2026-08-11 11:45:21 +08:00
chanx
cad2833c16 fix: Remove unused constant files (#18069)
### Summary

fix: Remove unused constant files
2026-08-11 11:16:06 +08:00
Jin Hai
cb3de6ae98 TS: add apache license (#18064)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 11:15:45 +08:00
Jin Hai
8379836179 Go: add context to test (#18070)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 11:12:34 +08:00
Jack
d6f6b6231f Fix(parser): don't collapse markdown doc into one item when a table is present (#18014) 2026-08-11 10:21:42 +08:00
Jack
dc73163908 Fix(tokenizer): pin alignment guards (#18011) 2026-08-11 10:09:07 +08:00
Jin Hai
e551167697 TS: fix comments (#18067)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-11 10:00:54 +08:00