Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
Open-weight coding agents reached up to 87.9% average task completion on a local data-prep benchmark.
The paper introduces an open-source evaluation framework for agents writing R code to prepare longitudinal population-study data. Its benchmark covers 20 tasks, including category harmonization and multi-wave merging, producing 102 variables from six sweeps of a British cohort study. The authors frame local open-weight models as a governance-friendly option where personal data cannot be sent to cloud services. State-of-the-art 31-35B parameter models came close to saturating the benchmark on consumer-grade hardware. ArXiv · AI/CL/LG's note
The paper introduces an open-source evaluation framework for agents writing R code to prepare longitudinal population-study data. Its benchmark covers 20 tasks, including category harmonization and multi-wave merging, producing 102 variables from six sweeps of a British cohort study. The authors frame local open-weight models as a governance-friendly option where personal data cannot be sent to cloud services. State-of-the-art 31-35B parameter models came close to saturating the benchmark on consumer-grade hardware. ArXiv · AI/CL/LG's note
score 4