TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
TACO cuts optimizer state for full-parameter LLM fine-tuning to near-zero levels while keeping AdamW-like accuracy in the paper’s tests.
The authors report OPT-13B persistent optimizer state falling from 27.7 GB with AdamW8bit to 0.16 GB with TACO. Peak training memory drops from 80.6 GB to 27.5 GB, with comparable accuracy and runtime. The method keeps first-order gradients but stores only low-precision gradient components per column. The paper says this makes full-parameter fine-tuning of 30-32B models possible on a single 80 GB H100 across several model families and tasks. ArXiv · AI/CL/LG's note
The authors report OPT-13B persistent optimizer state falling from 27.7 GB with AdamW8bit to 0.16 GB with TACO. Peak training memory drops from 80.6 GB to 27.5 GB, with comparable accuracy and runtime. The method keeps first-order gradients but stores only low-precision gradient components per column. The paper says this makes full-parameter fine-tuning of 30-32B models possible on a single 80 GB H100 across several model families and tasks. ArXiv · AI/CL/LG's note
score 4