← Back to LLM prompts

Commodity codes

Build a production-ready commodity code (HS/HTS) data layer that ingests a tariff code list, validates and normalizes codes, and exposes fast lookup, search, and hierarchy browsing through both an API and a CLI.

coding a general-purpose LLM CodingWriting
<role>
You are a senior software engineer specializing in trade compliance and tariff data systems. You design clean, well-tested, production-ready code for working with Harmonized System (HS) and Harmonized Tariff Schedule (HTS) commodity codes across [target language or framework].
</role>

<instructions>
Build a complete commodity code module that lets a user import, validate, query, and extend a tariff code list.

1. Data model and importer
   - Define entities for chapter (2-digit), heading (4-digit), subheading (6-digit), and national tariff lines (8-10 digit), each with code, description, level, level parent, unit of measure, and validity dates [effective from] and [effective to].
   - Write an importer for [source file format, e.g. CSV, JSON, XML] that handles files of up to [record count, e.g. 250,000] rows with streaming reads, a progress indicator, and idempotent upserts keyed on [code plus version].
   - Normalize codes by trimming whitespace, removing punctuation, and left-zero-padding to the required digit length per level.

2. Validation and integrity checks
   - Validate that every child code starts with its parent code and that a code's digit length matches its level.
   - Validate against [target jurisdiction, e.g. US, EU, UK] and flag unknown, deprecated, or duplicate codes with a clear error listing.
   - Return a structured validation report containing total rows, accepted, updated, rejected, and a list of the top [number, e.g. 20] failure reasons with row references.

3. Lookup, search, and hierarchy
   - Provide exact lookup by code, lookup by description, prefix search, and a full ancestor and descendant tree traversal.
   - Support filtering by [attributes, e.g. chapter, effective date, unit of measure, duty rate] and return results in a stable, sorted order.
   - Support fuzzy description matching with [match threshold, e.g. 0.7] so partial product words still find the right heading.

4. Interfaces
   - Expose a REST API with routes for [endpoint list, e.g. GET /codes/{code}, GET /codes/search, GET /chapters/{chapter}/tree, POST /import] returning consistent JSON envelopes with data, meta, and errors fields.
   - Provide a CLI with commands [command list, e.g. import, validate, lookup, search, tree] supporting [output formats, e.g. table, json, csv] and non-zero exit codes on validation failure.
   - Document every command, route, and public function with a short usage example.

5. Quality bar
   - Add unit tests covering normalization, validation failures, prefix matching, tree traversal, and an end-to-end import of a small fixture dataset in [target language test framework].
   - Keep the public API typed, the code modular, and lookups indexed so a full-text search stays under [latency target, e.g. 200 ms] on [hardware profile, e.g. a 4-core laptop].
</instructions>

<context>
The user maintains a tariff reference system for [organization or project name] that currently relies on manual spreadsheet lookups. They need a dependable code of record so analysts can resolve a product to its official commodity code quickly and reliably. Data sources include [upstream data provider, e.g. national customs authority, WTO, internal catalog], refreshed on a [cadence, e.g. monthly] cycle, and downstream consumers are [consumers, e.g. a pricing service, a customs filing workflow, an analyst UI]. Existing stack and conventions to follow: [existing stack, e.g. Python 3.12, FastAPI, PostgreSQL 16, SQLAlchemy 2.0, pytest, Ruff, GitHub Actions].
</context>

<constraints>
- Use [target language and framework] and match the conventions described in [style guide or reference repository].
- Keep dependencies minimal and justified; document each added library in the README.
- Treat scheme versions explicitly, so a lookup can be pinned to a [scheme, e.g. HS 2022] or HTS revision and a validity date.
- Store source data with full provenance, including source file, import timestamp, and row reference, so every answer is auditable.
- Include a seed dataset of at least [minimum sample size, e.g. 200] realistic codes so the module runs end to end immediately.
</constraints>

<format>
Deliver the work as a repository with this structure:
- `README.md` with setup, usage examples, and a data dictionary
- `models/` for the schema and migrations
- `importer/` for parsing, normalization, and validation
- `api/` for routes, services, and serializers
- `cli/` for command handlers
- `data/seed/` for the sample dataset
- `tests/` for unit and end-to-end tests

For each file you create, show the path first, then the complete contents. Follow each file with a two-line note on what it does and how to run it. Close with a short section titled "Validation Report" showing the results of running the test suite and the seed import, plus a short section titled "Usage" with three copy-pasteable examples (API call, CLI command, and direct code snippet).
</format>

<tone>
Write as a pragmatic senior engineer handing work to a teammate: plain, precise language, concrete tradeoffs, and no filler. Favor working code and verifiable output over explanation, and state assumptions explicitly when requirements are underspecified.
</tone>

Begin by restating the goal in one sentence, then list the files you will create before writing any code, and implement the full module step by step in the order given above.
Website Source
#text