VisionHOPE: Visual Backbones as Self-Modifying Learning Systems
VisionHOPE treats a vision backbone as a system whose memory and learning rules change together while it scans an image.
The paper builds on Nested Learning with five coupled memories for content, key/value generation, learning rate, and retention. The authors say an unconstrained self-referential update is unstable, so they add step-size controls and prove the scan dynamics are non-expansive. For 2D feature maps, VisionHOPE scans row and column chunks in four directions. They report competitive results on ImageNet-1K, COCO, and ADE20K. HF Daily Papers' note
The paper builds on Nested Learning with five coupled memories for content, key/value generation, learning rate, and retention. The authors say an unconstrained self-referential update is unstable, so they add step-size controls and prove the scan dynamics are non-expansive. For 2D feature maps, VisionHOPE scans row and column chunks in four directions. They report competitive results on ImageNet-1K, COCO, and ADE20K. HF Daily Papers' note
score 5