Megadose AI progress, ranked and analyzed.

RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs

· HF Daily Papers ·
A 53M-parameter model predicts free-text visual relationships from supplied regions in about 20 ms per frame.

RelateAnything takes an image, region boxes from any source, and a predicate vocabulary provided at inference as strings. The paper says it avoids object labels as inputs, using text embeddings for relations instead of a fixed learned predicate classifier. The author also introduces RA-4M, with 474k images and 4.3M verified relations, plus OV-SGG-Bench for cross-dataset evaluation. Reported results show 2.3-3.5x mean recall over the strongest comparable open-vocabulary method, and better scores than a 3B-parameter VLM scene-graph model at under 2% of its size. HF Daily Papers' note

score 5

Categories: Research