Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports
The paper releases two QA datasets built from 906 Panasonic technical documents, plus the pipeline used to make them.
The pipeline extracts layout-aware content, retrieves supporting evidence, and generates multiple-choice questions across five retrieval and answer-support settings. After filtering, each dataset has about 13,600 QA pairs with source documents and a held-out benchmark split. Fine-tuning sub-10B open models on the data raises Panasonic benchmark Set-Match Accuracy from 28.5% to 42.0% and F1 from 46.6% to 63.5%. The Claude-Opus-4.6-generated version is reported as cleaner and more effective than the Qwen3-generated version, but at roughly 100 times the cost. HF Daily Papers' note
The pipeline extracts layout-aware content, retrieves supporting evidence, and generates multiple-choice questions across five retrieval and answer-support settings. After filtering, each dataset has about 13,600 QA pairs with source documents and a held-out benchmark split. Fine-tuning sub-10B open models on the data raises Panasonic benchmark Set-Match Accuracy from 28.5% to 42.0% and F1 from 46.6% to 63.5%. The Claude-Opus-4.6-generated version is reported as cleaner and more effective than the Qwen3-generated version, but at roughly 100 times the cost. HF Daily Papers' note
score 4