Commit Graph

8515 Commits

Author SHA1 Message Date
Haruko386
1e46cef2d5 refactor: simplify sync task service and harden scheduler recovery (#18550)
### Summary

As title
2026-08-20 19:11:23 +08:00
mkaaad
cab70fdac3 feat(syncer): add Moodle connector with checkpoint resume (#18554)
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-08-20 19:05:16 +08:00
euvre
cc28035881 fix(ingestion): show data pipeline column for builtin pipeline ingestion logs (#18484) 2026-08-20 19:04:12 +08:00
Wang Qi
9357f79eb4 Fix: go switch user from superuser to normal user error (#18537) 2026-08-20 18:03:30 +08:00
euvre
f9120a0cff fix(compilation-template): emit wiki presets with the contracted "example" key (#18549) 2026-08-20 17:44:50 +08:00
Wang Qi
76bff4c915 Fix: switch image/ppt/email/audio from General to pipeline, or switch back, it report error (#18555) 2026-08-20 17:34:51 +08:00
Lynn
f7a2848dd5 Fix: ensure tenant_ocr_id is added (#18560) 2026-08-20 17:11:47 +08:00
buua436
a3df588951 feat: improve incremental Wiki compilation (#18557) 2026-08-20 16:45:34 +08:00
Sevenzuo
1f0285ee9b fix: preserve blocking GaussDB advisory locks (#18506) 2026-08-20 16:39:43 +08:00
Christian Uhl
f806001b3f fix(db): make negative GET_LOCK timeouts work on MariaDB (#18529)
### Summary

The fix turns a negative timeout into a finite 60 second wait inside
`MysqlDatabaseLock`, which both servers accept. Callers keep expressing
"block until available", and the equivalent PostgreSQL and GaussDB work
on the same `-1` semantics (#16346, #18506) is untouched. Non-negative
timeouts keep their current behaviour.

---------

Co-authored-by: Krim <git@krim.dev>
2026-08-20 16:33:46 +08:00
Haruko386
baaa6554e8 fix: get two VLM model type when list all models (#18546)
### Summary

As title
2026-08-20 16:32:49 +08:00
Jack
4eec5e4f35 feat(chunker): enforce a hard token cap in the TokenChunker merge (#18525) 2026-08-20 16:28:20 +08:00
mkaaad
0b7ec8a916 feat(syncer): add R2 connector with checkpoint resume (#18548)
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: Jin Hai <haijin.chn@gmail.com>
2026-08-20 16:26:49 +08:00
Christian Uhl
89d35d1b17 fix(auth): strip trailing slash from OIDC issuer before discovery (#18530)
### Summary

OIDC discovery builds its URL as
`f"{issuer}/.well-known/openid-configuration"`. Providers whose issuer
carries a trailing slash therefore get asked for a URL with a double
slash in it, which 404s. authentik is one of them: its issuer is
`https://auth.example.com/application/o/<app>/`, so the request goes to
`.../o/<app>//.well-known/openid-configuration` and login fails right at
the start with `Failed to fetch OIDC metadata`.

Stripping trailing slashes off the issuer before joining the well-known
path is enough. The `issuer` used later for ID token validation still
comes from the discovery document itself, so nothing else about the flow
changes.

Co-authored-by: Krim <git@krim.dev>
2026-08-20 16:25:29 +08:00
primorLee
3838c626aa fix(memory): make default prompts deterministic (#18547)
### Summary

Multi-type memory prompts were assembled from sets, so their instruction
and output sections could change order across Python processes with
different hash seeds. Because the generated default prompt is persisted
and later compared as a string, a restart could make an untouched
default look custom and prevent it from being regenerated when memory
types change.
2026-08-20 16:19:27 +08:00
Wangshu
77884f4a7c test: pin bare tenant_model.id resolution in resolve_model_config (#18528)
## Summary
- Adds a regression test for #18398

Fixes #18398

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-20 16:17:30 +08:00
YanZhang
a4f754a472 docs: add dataset docs & fix broken links in docs (#18553)
add dataset docs
fix broken links in docs
2026-08-20 16:12:23 +08:00
Haruko386
a88866f6b8 feat[syncer]: implement dingtalk AI table data source (#18539)
### Summary

As title
2026-08-20 16:10:08 +08:00
Haruko386
00ebe05dc5 fix: data source return two error once it occured (#18535)
### Summary

As title
2026-08-20 16:09:10 +08:00
Jack
20ee2dad22 fix(deepdoc): drop nested OCR boxes before table cell fill + expose 13/14 parity gaps (#18524) 2026-08-20 15:57:58 +08:00
Jack
d76c15e2a5 feat(title-chunker): enforce chunk_token_cap on the Go title-family chunks (#18455) (#18551) 2026-08-20 15:41:48 +08:00
jay77721
2022efa9fc refactor(ingestion): unify auto-metadata on modular metadata and hard-delete legacy flats (BuiltInMetadata carried, not LLM) (#18511) 2026-08-20 15:35:51 +08:00
Lynn
618c4599b1 Fix: handle tenant_llm without id (#18541) 2026-08-20 14:12:52 +08:00
Loi Nguyen
e8aeddc5f7 fix(dataset): preserve disabled RAPTOR and GraphRAG settings (#17665)
### Summary

Preserve explicitly saved `false` values for RAPTOR and GraphRAG when
hydrating the dataset Configuration form. The existing form defaults
still apply when either enable flag is absent.

Adds hook-level regression coverage for both explicit disablement and
default fallback behavior.

Fixes #17654
2026-08-20 13:46:28 +08:00
iwasaki
9d46a70fc0 Fix: respect IME composition state on Enter key in dataset creating dialog (#17150)
## What

Pressing Enter to confirm IME (e.g. Japanese) text conversion in the
knowledge base name field of the "Create knowledge base" dialog was
incorrectly treated as the dialog's submit trigger. The
composition-confirm Enter both let the IME finish composing and
triggered the dialog's Enter handler, which called `preventDefault()`
and `form.requestSubmit()` — causing input like "アルゴ" to be duplicated
as "アルゴアルゴ".
2026-08-20 13:45:09 +08:00
Brian Sparker
09a9b629ee feat(chat): add You.com web search provider (#18478)
### Summary

Adds You.com as a built-in Web Search provider for RAGFlow Chat,
alongside Tavily and Querit, using the provider-neutral dispatch #17813
put in place. No changes to existing Tavily or Querit behaviour.

You.com runs its own web index and returns several extracted passages
per result rather than a single meta description, so retrieved chunks
arrive with usable context.

---------

Co-authored-by: Brian Sparker <brainsparker@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 12:12:18 +08:00
Jin Hai
9421796ccf Update release notes (#18544)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-20 12:05:57 +08:00
Jack
84bed4dec5 refactor(chunker): replace QAChunker/PresentationChunker with PairChunker/PageChunker, drop TagChunker (#18523) 2026-08-20 11:53:00 +08:00
I Kartik Reddy
ab40c90118 fix: avoid quadratic dedup in RAG merge paths (#18135)
## Summary
Fixes #18025. Both merge paths deduplicated IDs by scanning a plain list
(`item not in list`) inside a loop while appending — O(n²) per merge.
Replaced with a set-backed `seen` check alongside the existing ordered
list: same order, same dedup result, O(n).

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:50:27 +08:00
Haruko386
994148278d fix: ollama chat error (#18540) 2026-08-20 11:16:14 +08:00
Rootkit
834dfc1906 Add ulimits for cpu and gpu services in Docker Compose (#17205)
### Summary

fix OSError: [Errno 24] Too many open files
2026-08-20 10:43:22 +08:00
krishna soni
17a55558b7 feat(cli): add --version flag to cli (#18534)
### Summary

This PR adds a `--version` (and `-V`) flag to the `ragflow-cli` tool.

### Usage

```bash
ragflow-cli --version
2026-08-20 10:33:03 +08:00
Loi Nguyen
f420e47f95 fix(ollama): preserve multimodal images (#17663)
### Summary

- Convert multimodal text blocks into Ollama message content.
- Attach image references through Ollama native images arrays for
synchronous and streaming chat requests.
- Strip data-URI headers while preserving raw base64 and URL image
values.
- Add regression coverage for both request paths.

Closes #17332

---------

Co-authored-by: Codex <codex@openai.com>
2026-08-20 10:15:15 +08:00
Jack
c90a0b7a87 feat(deepdoc): Python-intermediate replay parity harness + table cell-fill fixes (#18516) dev-20260820 2026-08-19 21:45:45 +08:00
euvre
e5a8b10120 fix(chat): drain deep-research goroutines before closing the async channel (#18504) 2026-08-19 21:23:51 +08:00
Dhruv Diwakirti
86c25068fa fix: keep OCR text when no image2text model is configured (#18012)
An image whose OCR text is shorter than the CV LLM threshold produces zero chunks when the tenant has no image2text model configured. The extracted text is discarded.
2026-08-19 21:20:32 +08:00
Jin Hai
057cf41415 TS: update thinking description (#18521)
Signed-off-by: Jin Hai <haijin.chn@gmail.com>
2026-08-19 20:06:30 +08:00
Linpeng cheng
a9323ca554 fix: pass Infinity vector similarity weight (#17453)
## Summary

- Pass `vector_similarity_weight` from Python and Go retrieval requests
into Infinity's weighted fusion expression.
- Keep fusion weights ordered as text first and vector second, with the
existing default vector weight of `0.3`.

---------

Co-authored-by: chenglinpeng <1042527908@qq.com>
2026-08-19 20:06:02 +08:00
mkaaad
a3170af846 Feat: connector go imap (#18513)
add imap connector to go

---------

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-08-19 19:23:26 +08:00
lawrence
c917d8b4e3 fix: clear stale tenant model ids (#18210)
### Summary

Clear the paired tenant model ID when a model selection is explicitly
cleared in a whitelisted API request.

Replaces #18205.

Co-authored-by: Jin Hai <haijin.chn@gmail.com>
2026-08-19 19:05:03 +08:00
Serply
371d83c2d8 feat(chat): add Serply web search provider (#18475)
### Summary

This PR adds [Serply](https://serply.io) as a third web search provider
for chat assistants, alongside the existing Tavily and Querit options.
2026-08-19 18:59:01 +08:00
peewee92
93b38808e6 feat(chat): self-check contradictory prompt/dataset settings (#5703) (#18274)
## Summary

Closes #5703.

Users who delete the hardcoded `{knowledge}` placeholder from the system
prompt while datasets are selected can still retrieve the right chunks,
but the assistant answers as if nothing were found — because the
retrieved content has nowhere to be injected. Likewise, a non-empty
*empty response* with **no** dataset selected fires on every turn
(nothing can ever be retrieved). This PR adds a save-time self-check
that prompts the user about both contradictory configurations, as
requested in the issue.

Co-authored-by: peewee92 <20059253+peewee92@users.noreply.github.com>
2026-08-19 18:56:50 +08:00
Haruko386
86ecdd8588 fix: chat did not pass model config to provider (#18485)
### Summary

As title
2026-08-19 18:43:15 +08:00
Loong
86c520a336 fix(nlp): differentiate alphabetic OOV term weights (#18470)
### Summary

Closes #18414.

`rag/res/term.freq` is not shipped, and both term-weight implementations
therefore assigned the same `300` fallback frequency to every lowercase
Latin token. With no tokenizer frequency, NER, or POS signal, function
words and content words received identical lexical boosts.

This PR adds the same bounded out-of-vocabulary prior to Python and Go:

- Use it only when the explicit DF dictionary or tokenizer has no
frequency.
- Count Latin, Greek, and Cyrillic letters, including uppercase and
accented forms.
- Keep the existing frequency of `300` for words up to three letters,
halve it every two additional letters, and clamp it at `10`.
- Reject digits, underscores, and logographic terms so Chinese and other
existing fine-grained-tokenizer paths are unchanged.
- Treat an absent optional `term.freq` as the supported fallback path
without a startup warning, while still logging inaccessible or malformed
dictionaries.

A corpus-derived table was intentionally not added: that would require
provenance/licensing decisions, language detection, and handling
cross-language homographs. The bounded prior is deterministic,
dependency-free, and fixes the equal-weight degradation for
whitespace-delimited alphabetic languages without claiming
corpus-specific precision.

Python and Go consume one shared fixture covering ASCII, uppercase,
accented Latin, Greek, Cyrillic, separators, invalid mixed tokens, and a
CJK non-match. Both sides also verify the issue's ordering (`was <
largest < supplier < equipment`) and that an explicit dictionary entry
still takes precedence.


Co-authored-by: Loong <184861530+yzl0ng@users.noreply.github.com>
2026-08-19 18:37:04 +08:00
Loong
a6b5e985c4 fix(memory): require semantic valid_at timestamp (#18462)
## Summary

- require a `valid_at` timestamp in the semantic-memory output schema
- tell the extraction model to use conversation time when a fact has no
date of its own
- add a regression test for the assembled semantic prompt

This addresses the deterministic prompt inconsistency reported in
#18415. The invalid-timestamp fallback is intentionally left unchanged
because selecting its replacement policy requires a separate design
decision.
2026-08-19 18:32:27 +08:00
Loong
1e147e0c0b fix(memory): normalize invalid extraction timestamps (#18463)
## Summary

- allow ISO 8601 normalization callers to provide an explicit fallback
while preserving the existing default behavior
- use the extraction conversation time when `valid_at` is missing or
invalid
- clear an invalid optional `invalid_at` instead of writing an
unparseable value
- include the rejected timestamp value in the error log

This addresses the timestamp write-through portion of #18415. It is
intentionally separate from #18462, which fixes the semantic output
prompt.

Co-authored-by: Loong <184861530+yzl0ng@users.noreply.github.com>
2026-08-19 18:31:51 +08:00
Loong
bf9f06c566 fix(memory): honor custom extraction prompts (#18461)
### Summary

- Forward each memory's stored system_prompt and user_prompt to
extract_by_llm.
- Cover both immediate save and queued extraction paths with focused
regression tests.
- Preserve the existing default-prompt fallback when stored prompts are
empty.

Fixes #18413.
2026-08-19 18:28:36 +08:00
Sevenzuo
dd1f335ba2 fix: preserve extensionless document suffix on GaussDB (#18483)
### Summary

RAGFlow's "Create empty document" flow accepts names without a file
extension. The `POST /datasets/<dataset_id>/documents?type=empty` route
calls `_upload_empty_document()`, where `Path(name).suffix.lstrip(".")`
returns `""`.

In GaussDB's A/ORA compatibility mode, that empty string is persisted as
SQL `NULL`. Because `document.suffix` was defined as `NOT NULL`, the
insert failed with a constraint violation.
2026-08-19 18:26:14 +08:00
Ali Farhan
ece9638f94 fix(api): report why provider model discovery failed instead of swallowing it (#18027)
### Summary

Providers whose static catalogue is empty discover their models by
calling the base URL the user typed. `verify_api_key` wrapped that call
in a bare `except Exception: pass` and then returned a flat `No models
found for provider 'X'`, so an unreachable host, a closed port, a wrong
scheme and a bad TLS setup all produced the same sentence, with the
actual error discarded and not even logged.
2026-08-19 18:24:10 +08:00
Haruko386
490afa46e4 refactor: remove some unused model types (#18502)
### Summary

As title
2026-08-19 18:15:43 +08:00