Abstract:
One of the biggest challenges when it comes to retrieving information from diversified and complex datasets, such as PDFs, is their unstructured nature and the possibility of re ceiving irrelevant or incorrect responses. The project focuses on building a conversational AI system that can efficiently extract information from PDFs and offer responses in both English and Hindi. Integration with WhatsApp enhances real-time interaction for users, thereby increasing the usability and engagement of the system, making it versatile and intuitive. The system uses the Llama-3.1-70b-versatile model of Chat GROQ, along with a FAISS vector store for document querying. The system can handle vast datasets across multiple PDFs with great accuracy in the responses it provides. LangChain allows the system to ensure a coherent flow of conversation and interactions in real time. It is designed to enable easy handling of bilingual queries. The multi-platform integra tion ensures robust yet user-friendly performance on the web and messaging platforms. In essence, the use of such advanced techniques for retrieval and including traceability to sources enables it to avoid the challenges of big and complex document sets and present a reliable result in the domain of health information retrieval.