Langhuan can now process PDFs through an async parsing pipeline and archive parsed assets like images to local or S3-compatible storage. A PDF is no longer a one-shot black-box upload — it leaves behind an intermediate structure you can keep indexing and tracing.
PDF parsing doesn’t block the request
The upload endpoint only receives the document, records its status, and enqueues a background task. Once MinerU Cloud finishes parsing, workers continue processing the result, and the frontend can see the current stage through the task status.
- The parsing provider plugs in through an independent capability contract, so credentials don’t get scattered through business flows.
- Parse failures, task retries, and completion state all live in the document lifecycle.
- Async tasks for the same document can safely re-run, avoiding duplicate asset or index writes.
Assets are part of the knowledge, too
Images, table-related files, and relative-path resources from the PDF go into the AssetStore instead of being left in a temp directory. Local and S3-compatible storage implementations share the same asset-resolution boundary, so you can switch by environment at deploy time.
From parsing to searchable
After parsing completes, the system continues with asset archiving, content chunking, and index publishing. Each step preserves the link between document and asset, so a retrieval result can return to the original document rather than leaving behind text with no provenance.