Simple RAG System Builder
coding a general-purpose LLM Customer SupportCoding
<role>You are an expert AI engineer specializing in building production-ready RAG (Retrieval-Augmented Generation) systems using modern LLMs, vector databases, and embedding models.</role> <task>Build a complete, modular RAG pipeline that ingests documents, creates searchable vector embeddings, and answers user queries with grounded, cited responses.</task> <context> <project_name>[project name]</project_name> <document_sources>[list of document paths, URLs, or directories to ingest]</document_sources> <embedding_model>[embedding model name, e.g., text-embedding-3-small, bge-small-en-v1.5]</embedding_model> <llm_model>[LLM for generation, e.g., gpt-4o-mini, llama-3.1-8b-instruct]</llm_model> <vector_db>[vector database choice, e.g., chroma, faiss, pinecone, weaviate]</vector_db> <chunk_size>[chunk size in tokens, e.g., 512]</chunk_size> <chunk_overlap>[chunk overlap in tokens, e.g., 50]</chunk_overlap> <top_k>[number of chunks to retrieve, e.g., 4]</top_k> <language>[response language, e.g., English]</language> </context> <constraints> - Use only well-maintained, popular libraries (langchain, llama-index, or native SDKs) - Implement proper error handling and logging throughout - Include document metadata preservation (source, page, date, etc.) - Support multiple file formats (PDF, TXT, MD, HTML, DOCX) - Implement hybrid search (vector + keyword) for better recall - Add citation tracking so every answer references source chunks - Include a simple CLI or API interface for querying - Write clean, typed, documented Python code - Provide a requirements.txt and README with usage examples - No hardcoded secrets — use environment variables </constraints> <format> Deliver a complete project structure: ``` [project_name]/ ├── src/ │ ├── __init__.py │ ├── config.py # Configuration management │ ├── ingestion.py # Document loading & chunking │ ├── embeddings.py # Embedding model wrapper │ ├── vectorstore.py # Vector DB operations │ ├── retrieval.py # Hybrid search & reranking │ ├── generation.py # LLM response with citations │ ├── pipeline.py # Main RAG pipeline orchestration │ └── cli.py # Command-line interface ├── tests/ │ └── test_pipeline.py ├── data/ # Place documents here ├── .env.example ├── requirements.txt ├── README.md └── main.py # Entry point ``` Each module must include: - Type hints for all functions - Docstrings explaining purpose, args, returns - Structured logging (structlog or standard logging) - Unit-testable design (dependency injection) </format> <tone>Professional, pragmatic, and code-focused. Prioritize clarity, maintainability, and correctness over cleverness. Explain architectural decisions in comments where non-obvious.</tone> <final_instruction>Generate the complete project code now. Start with config.py and requirements.txt, then build each module in dependency order. Ensure the final pipeline can be run with `python main.py --query "[sample question]"` and returns a cited answer.</final_instruction>
#text