Abstract: This paper addresses shortcomings in the current combined application of large language models (LLMs) and vector knowledge bases by proposing a novel method for text information extraction. The method integrates the powerful language understanding capabilities of LLMs with the robust matching ability of two-stage information filtering, aiming to improve both the hit rate for retrieving relevant background knowledge and the accuracy of text information extraction. Experimental results on a real-world accident investigation report dataset demon⁃ strate that the proposed method achieves a significant 37% improvement in extraction accuracy compared to baseline methods. Additionally, comparative results across data from different industries further validate the method’s strong generalization performance, suggesting its poten⁃ tial as an effective technical solution for other long-text information extraction tasks.
Keywords: large language models; two-stage information filtering; text information extraction