← Back to LLM prompts

HyDE-Based Post Retrieval System

A research-focused prompt that implements the Hypothetical Document Embeddings (HyDE) technique to enhance post retrieval accuracy by generating synthetic relevant documents for improved semantic search and ranking.

research a general-purpose LLM AnalysisPrompt Engineering
<role>You are an expert information retrieval researcher specializing in dense retrieval augmentation techniques, particularly Hypothetical Document Embeddings (HyDE) for improving semantic search quality.</role>

<task>Design and implement a HyDE-enhanced post retrieval pipeline that generates hypothetical relevant posts for a given query, embeds them alongside real posts, and retrieves the most semantically relevant actual posts from a target corpus.</task>

<context>
- Target domain: [target domain or platform, e.g., technical forums, social media, academic discussions]
- Corpus size: [approximate number of posts in the retrieval corpus]
- Query types: [typical user query patterns, e.g., natural language questions, keyword searches, conversational queries]
- Embedding model: [preferred embedding model, e.g., text-embedding-3-large, bge-large-en-v1.5, instructor-xl]
- Generation model: [LLM for hypothetical document generation, e.g., GPT-4, Claude-3, Llama-3-70B]
- Evaluation metrics: [retrieval metrics to optimize, e.g., nDCG@10, Recall@100, MRR]
- Baseline system: [current retrieval method to compare against, e.g., BM25, dense retrieval only, hybrid search]
</context>

<constraints>
- Generate exactly [number of hypothetical documents, e.g., 5-10] hypothetical posts per query
- Ensure hypothetical posts are plausible but clearly marked as synthetic
- Maintain query intent fidelity in generated hypothetical documents
- Use consistent embedding dimensions across hypothetical and real documents
- Implement efficient batching for hypothetical document generation
- Include ablation study design comparing: (1) dense only, (2) HyDE only, (3) hybrid HyDE + dense
- Document all hyperparameters: temperature, top-p, max tokens for generation
- Ensure reproducibility with fixed random seeds
</constraints>

<format>
Provide a structured research implementation plan with:

1. **HyDE Pipeline Architecture**
   - Query preprocessing steps
   - Hypothetical document generation prompt template
   - Embedding strategy (separate vs. joint embedding space)
   - Retrieval and re-ranking logic

2. **Experimental Design**
   - Dataset splits (train/val/test) with statistics
   - Baseline configurations
   - Ablation variants
   - Statistical significance testing approach

3. **Implementation Specifications**
   - Pseudocode for core HyDE retrieval loop
   - Memory and compute requirements
   - Latency optimization strategies
   - Error handling for generation failures

4. **Evaluation Framework**
   - Metric computation code snippets
   - Result visualization specifications
   - Failure case analysis template

5. **Reproducibility Package**
   - Configuration file schema (YAML/JSON)
   - Dependency versions
   - Random seed documentation
   - Hardware specifications
</format>

<tone>Technical, precise, research-oriented, and methodical with emphasis on experimental rigor and reproducibility.</tone>

<final_instruction>Generate the complete HyDE post retrieval research implementation plan now, filling in all placeholders with realistic values for a [target domain] retrieval system evaluating on [benchmark dataset name] with [corpus size] posts.</final_instruction>
Website Source
#text