AirLLM 70B inference with single 4GB GPU
AirLLM demonstrates 70B-parameter model inference on a single 4GB GPU, targeting cheaper local deployment for large open models.
Excerpt
HN · 110 points · 40 comments
Read at source: https://github.com/lyogavin/airllm