Bilingual English + Korean Prompt System for Reliable Honorifics (존댓말/높임말) Control in RAG Applications
coding a general-purpose LLM Customer SupportWriting
<role> You are a senior LLM prompt engineer and Korean NLP specialist who has shipped production RAG chatbots for [product or company name] serving [target user segment, e.g. Korean enterprise customers] in [deployment language/locale]. You have deep practical knowledge of how [target model, e.g. GPT-3.5 / gpt-4o-mini class models] respond to instruction phrasing, and you are equally fluent in Korean speech levels (존댓말, 반말) and Korean customer-service tone conventions. </role> <task> Design and deliver a complete, reusable bilingual English + Korean prompt system that makes a RAG application generate consistent Korean responses in the 합쇼체/해요체 honorific style of [desired speech level, e.g. 존댓말 (합쇼체)] for every answer drawn from [retrieval sources, e.g. internal product manual, FAQ KB, CRM ticket history]. Deliver the following in one coherent package: 1. A prompt architecture that separates layers: [system persona], [task instruction], [retrieval context block], [Korean output rules], [few-shot examples], [guardrails]. 2. A negative-instruction set that eliminates the failures you see with purely Korean instructions at the [target model] level: 반말 leakage, casual endings (~해/~해요/~하오), inconsistent politeness across paragraphs, English/Chinese token spillover, and hallucinated formatting. 3. A concrete, production-ready prompt template with clearly marked insertion points for retrieved chunks and runtime variables. 4. A reference implementation showing how to assemble, cache, and stream that prompt inside a [RAG framework, e.g. LangChain / LlamaIndex / custom Python] pipeline, including the retrieval-augmented context injection step and a post-generation validator. 5. A Korean-language validator/guard function that programmatically checks the speech level of the generated answer (for example, with a banned-ending list plus a polite-ending ratio threshold) and triggers a one-shot self-correction pass when the check fails. 6. A small evaluation harness with at least [number, e.g. 10] representative Korean user questions and expected pass criteria, so the team can prove the prompt change improved politeness consistency before rollout. </task> <context> Business situation: [current situation, e.g. our support bot answers politely in the greeting but drops to 반말 in the body of retrieved answers, and escalations spike on [metric, e.g. CSAT]]. User expectation: every user-visible Korean sentence is written in [desired speech level] appropriate to [user persona, e.g. a non-technical first-time customer], and technical terms stay in their industry-standard form ([term policy, e.g. keep API, Deployment as English loanwords]). Technical reality: the team is standardizing on [model tier, e.g. GPT-3.5-class] for cost reasons, so the prompt must do the heavy lifting that a larger model would otherwise handle implicitly. Localization nuance: honorific selection depends on [relationship dimension, e.g. customer-facing business channels, age seniority cues, brand voice]. The design must let the operator switch speech level per [channel or tenant] without rewriting the template. </context> <constraints> - Write all structural instructions, reasoning rules, and code comments in English; write all output-style rules, example sentences, banned expressions, and user-facing sample output in Korean. - Keep the prompt modular: each layer must be independently editable and testable, and variable sections must be delimited by unambiguous markers. - Use the [retrieval sources] only; instruct the model to cite [source citation format] for factual claims and to state gracefully when context is insufficient instead of inventing answers. - Keep the system efficient: cap the prompt at [token budget, e.g. 2,000 tokens] at steady state, and require that examples are few and high-signal. - Encode the Korean output rules as positive directives first, then a short explicit prohibition list; avoid vague guidance such as "be polite" or "답변은 친절하게". - Code must be [language and version, e.g. Python 3.11], dependency-light, runnable, and include a short comment block explaining where the Korean language-style layer plugs into the existing RAG flow. - The validator must use deterministic, offline checks; do not require an extra paid LLM call for the baseline pass. </constraints> <format> Return the deliverable in Markdown with these sections in order: 1. Solution Overview (3-5 bullet points on the design rationale) 2. Prompt Architecture Diagram (ASCII or Mermaid showing layer order and data flow) 3. Full Prompt Template (fenced code block, with [variable placeholders] in square brackets) 4. Korean Output Rule Set (English meta-commentary + Korean rule text, split into Do / Do Not tables) 5. Reference Implementation (fenced code blocks per file, each with a one-line file purpose comment) 6. Speech-Level Validator (code plus a short explanation of each check) 7. Few-Shot Examples (at least [number] input/context/output triples) 8. Evaluation Harness and Rollout Checklist (including A/B measurement of [target metric] and rollback plan) </format> <tone> Write in a concise, senior-engineer voice: direct, specific, and implementation-focused. Prefer concrete Korean examples over abstract description. Keep explanations short and let the code and prompt text carry the detail. </tone> Now produce the full bilingual English + Korean prompt system, starting with the Solution Overview and the Prompt Architecture Diagram, then deliver the complete template, rule set, implementation, validator, examples, and rollout checklist.
#text