Text Preprocessing Pipeline
data a general-purpose LLM Customer SupportCoding
<role>You are an expert NLP engineer and data scientist specializing in text preprocessing for machine learning, analytics, and information retrieval systems.</role> <task>Build a complete, production-ready text preprocessing pipeline that transforms [raw text data] into clean, normalized, and structured output suitable for [downstream task: e.g., sentiment analysis, topic modeling, classification, search indexing].</task> <context> - Input: [raw text data] — can be a single string, list of documents, pandas Series, or path to [input file format: CSV, JSON, TXT, Parquet] - Target language(s): [language(s): e.g., English, Spanish, multilingual] - Domain: [domain: e.g., social media, legal documents, biomedical, customer reviews, news] - Downstream task: [downstream task] — determines which preprocessing steps are optimal - Volume: [data volume: e.g., <10K docs, 10K-1M, >1M] — affects performance considerations </context> <constraints> - Use only well-maintained, standard libraries (spaCy, NLTK, regex, pandas, scikit-learn, huggingface tokenizers) - Preserve document boundaries and metadata (document IDs, timestamps, labels) throughout - Handle edge cases gracefully: empty strings, encoding issues, mixed languages, emojis, URLs, mentions - Provide configurable toggles for each preprocessing step - Log processing statistics (tokens removed, documents affected, processing time) - Output must be reproducible with a fixed random seed where applicable - Memory-efficient: support streaming/batch processing for large datasets </constraints> <format> Return a Python class `TextPreprocessingPipeline` with: 1. `__init__(config: PreprocessingConfig)` — accepts a dataclass/config object with all toggles 2. `fit(documents: Iterable[str]) -> Self` — learns vocabulary, statistics (optional) 3. `transform(documents: Iterable[str]) -> List[ProcessedDocument]` — applies pipeline 4. `fit_transform(documents: Iterable[str]) -> List[ProcessedDocument]` — convenience method 5. `get_stats() -> Dict` — returns processing statistics `ProcessedDocument` = NamedTuple with fields: - `doc_id: str` - `original_text: str` - `cleaned_text: str` - `tokens: List[str]` - `lemmas: List[str]` - `pos_tags: List[str]` - `entities: List[Tuple[str, str]]` # (entity_text, entity_label) - `metadata: Dict` # preserved input metadata `PreprocessingConfig` toggles (all default True unless noted): - `lowercase: bool` - `remove_html: bool` - `remove_urls: bool` - `remove_emails: bool` - `remove_mentions: bool` # @username - `remove_hashtags: bool` # keep #topic or remove entirely - `expand_contractions: bool` # "don't" → "do not" - `normalize_unicode: bool` # NFKC normalization - `remove_accents: bool` # False for languages needing diacritics - `remove_numbers: bool` # False - `remove_punctuation: bool` # keep sentence punctuation option - `tokenizer: Literal["spacy", "nltk", "hf", "whitespace"]` - `remove_stopwords: bool` - `stopword_list: Optional[Set[str]]` # custom - `apply_lemmatization: bool` - `apply_stemming: bool` # mutually exclusive with lemmatization - `min_token_length: int` # default 2 - `max_token_length: int` # default 50 - `remove_custom_patterns: List[str]` # regex patterns - `preserve_case_for: List[str]` # entity types to keep cased - `language: str` # ISO 639-1 code - `batch_size: int` # for streaming - `n_jobs: int` # parallel processing - `random_seed: int` </format> <tone>Professional, precise, and implementation-focused. Code must be production-quality with type hints, docstrings, and error handling.</tone> <final_instruction>Generate the complete Python implementation of `TextPreprocessingPipeline` with `PreprocessingConfig` dataclass, `ProcessedDocument` NamedTuple, and a usage example demonstrating configuration for [downstream task] on [domain] data in [language(s)].</final_instruction>
#text