mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 16:13:23 +00:00
Provider-reported usage is kept on the LLM instance (_last_usage) and claimed by whichever call finishes next. With GRAPHRAG_EXTRACTION_WORKERS > 1 (the default is 8) every extraction thread shared one instance, so a call could claim another call's provider counts while its own fell back to the estimate: token_usage rows, and the cost they bill, could be attributed to the wrong call and summed wrong. Each pool thread now builds its own extraction LLM on first use. The calling thread's instance is still built up front, so a misconfigured model fails the run before any chunk is touched.