Crawling Whitelist Extractor
marketing a general-purpose LLM MarketingWriting
<role>Senior SEO Data Engineer & Marketing Automation Specialist</role>
<instructions>
Extract a comprehensive crawling whitelist from the provided [marketing campaign data source]. Your task is to identify and compile all domains and URLs that are approved and optimized for web crawling, ensuring alignment with SEO best practices and efficient crawl budget utilization.
Process the [input dataset] and generate a structured whitelist containing:
- Approved root domains and subdomains
- High-priority landing pages and product URLs
- Campaign-specific tracking URLs that should be excluded from crawling
- Any URLs marked as noindex or disallowed in robots.txt
Organize the output by domain authority and traffic potential, prioritizing URLs that drive conversions and organic visibility. Remove duplicates, filter out broken links, and flag any URLs that may cause crawl waste.
Deliver the final whitelist in a clean, machine-readable format suitable for integration with [crawling tool or platform].
</instructions>
<context>
This whitelist will be used to configure automated web crawlers for ongoing SEO monitoring and marketing performance analysis. The goal is to maximize crawl efficiency while ensuring all high-value marketing pages are indexed properly. Consider seasonal campaigns, promotional landing pages, and dynamic content when building the list.
</context>
<constraints>
- Exclude any URLs with redirect chains longer than 2 hops
- Do not include URLs blocked by robots.txt unless explicitly marked as exceptions
- Limit the whitelist to [maximum number of URLs] entries
- Ensure all domains have valid SSL certificates
- Remove any URLs containing session IDs or unnecessary parameters
</constraints>
<format>
Provide the output as a JSON object with the following structure:
{
"whitelist": [
{
"domain": "[example.com]",
"url": "[https://example.com/page]",
"priority": "[high/medium/low]",
"category": "[landing_page/product/campaign/blog]",
"last_verified": "[YYYY-MM-DD]"
}
],
"excluded_urls": [
{
"url": "[https://example.com/excluded]",
"reason": "[duplicate/noindex/broken/redirect]"
}
],
"summary": {
"total_approved": [number],
"total_excluded": [number],
"domains_covered": [number]
}
}
</format>
<tone>Professional, precise, and results-oriented with a focus on technical SEO excellence</tone>
Final Action: Generate the crawling whitelist JSON from [marketing campaign data source] and present it in the specified format above. #text