← Back to LLM prompts

Text Preprocessing Pipeline

A comprehensive, configurable text preprocessing pipeline that transforms raw text data into clean, analysis-ready format. Handles cleaning, normalization, tokenization, and linguistic processing with customizable options for different languages and use cases.

data a general-purpose LLM Customer SupportCoding
<role>You are an expert NLP engineer and data scientist specializing in text preprocessing for machine learning, analytics, and information retrieval systems.</role>

<task>Build a complete, production-ready text preprocessing pipeline that transforms [raw text data] into clean, normalized, and structured output suitable for [downstream task: e.g., sentiment analysis, topic modeling, classification, search indexing].</task>

<context>
- Input: [raw text data] — can be a single string, list of documents, pandas Series, or path to [input file format: CSV, JSON, TXT, Parquet]
- Target language(s): [language(s): e.g., English, Spanish, multilingual]
- Domain: [domain: e.g., social media, legal documents, biomedical, customer reviews, news]
- Downstream task: [downstream task] — determines which preprocessing steps are optimal
- Volume: [data volume: e.g., <10K docs, 10K-1M, >1M] — affects performance considerations
</context>

<constraints>
- Use only well-maintained, standard libraries (spaCy, NLTK, regex, pandas, scikit-learn, huggingface tokenizers)
- Preserve document boundaries and metadata (document IDs, timestamps, labels) throughout
- Handle edge cases gracefully: empty strings, encoding issues, mixed languages, emojis, URLs, mentions
- Provide configurable toggles for each preprocessing step
- Log processing statistics (tokens removed, documents affected, processing time)
- Output must be reproducible with a fixed random seed where applicable
- Memory-efficient: support streaming/batch processing for large datasets
</constraints>

<format>
Return a Python class `TextPreprocessingPipeline` with:
1. `__init__(config: PreprocessingConfig)` — accepts a dataclass/config object with all toggles
2. `fit(documents: Iterable[str]) -> Self` — learns vocabulary, statistics (optional)
3. `transform(documents: Iterable[str]) -> List[ProcessedDocument]` — applies pipeline
4. `fit_transform(documents: Iterable[str]) -> List[ProcessedDocument]` — convenience method
5. `get_stats() -> Dict` — returns processing statistics

`ProcessedDocument` = NamedTuple with fields:
- `doc_id: str`
- `original_text: str`
- `cleaned_text: str`
- `tokens: List[str]`
- `lemmas: List[str]`
- `pos_tags: List[str]`
- `entities: List[Tuple[str, str]]`  # (entity_text, entity_label)
- `metadata: Dict`  # preserved input metadata

`PreprocessingConfig` toggles (all default True unless noted):
- `lowercase: bool`
- `remove_html: bool`
- `remove_urls: bool`
- `remove_emails: bool`
- `remove_mentions: bool`  # @username
- `remove_hashtags: bool`  # keep #topic or remove entirely
- `expand_contractions: bool`  # "don't" → "do not"
- `normalize_unicode: bool`  # NFKC normalization
- `remove_accents: bool`  # False for languages needing diacritics
- `remove_numbers: bool`  # False
- `remove_punctuation: bool`  # keep sentence punctuation option
- `tokenizer: Literal["spacy", "nltk", "hf", "whitespace"]`
- `remove_stopwords: bool`
- `stopword_list: Optional[Set[str]]`  # custom
- `apply_lemmatization: bool`
- `apply_stemming: bool`  # mutually exclusive with lemmatization
- `min_token_length: int`  # default 2
- `max_token_length: int`  # default 50
- `remove_custom_patterns: List[str]`  # regex patterns
- `preserve_case_for: List[str]`  # entity types to keep cased
- `language: str`  # ISO 639-1 code
- `batch_size: int`  # for streaming
- `n_jobs: int`  # parallel processing
- `random_seed: int`
</format>

<tone>Professional, precise, and implementation-focused. Code must be production-quality with type hints, docstrings, and error handling.</tone>

<final_instruction>Generate the complete Python implementation of `TextPreprocessingPipeline` with `PreprocessingConfig` dataclass, `ProcessedDocument` NamedTuple, and a usage example demonstrating configuration for [downstream task] on [domain] data in [language(s)].</final_instruction>
Website Source
#text