Data Pipeline
The site follows a strict data pipeline to ensure accuracy and traceability:
- Source acquisition — Official PDF documents and structured data files are downloaded from Soumu and e-Stat. No third-party data sources are used.
- Extraction — The build scripts parse official PDFs and JSON files to extract the complete JSIC hierarchy: 20 sections, 98 major groups, 519 groups, and 1411 industry codes.
- Validation — Each extracted code is validated against the official structure. Parent-child relationships are verified. Missing or orphaned codes are flagged.
- Enrichment — Official annotations (overview, inclusions, exclusions, notes) are linked to each code. Crosswalks to ISIC, NACE, and NAICS are applied at the section level.
- Generation — The build script generates static HTML pages for every code, section, and classification level. Each page includes source references and hierarchy navigation.