AVA-Encoder: Towards Agent-Native Video Representation Learning
AVA-Encoder turns films into editable knowledge graphs for video agents, then reconstructs them back into video.
The paper frames that structure as a way for creative agents to learn from high-quality human films. Its graph stores text descriptions alongside linked image, audio, and video assets, with typed edges for relationships agents can query and edit. The authors say reconstruction differences are fed back as natural-language update directions for training and optional test-time refinement. They report a 20.7-point gain over the strongest external baseline and release the framework, benchmark, and film KG dataset. ArXiv · AI/CL/LG's note
The paper frames that structure as a way for creative agents to learn from high-quality human films. Its graph stores text descriptions alongside linked image, audio, and video assets, with typed edges for relationships agents can query and edit. The authors say reconstruction differences are fed back as natural-language update directions for training and optional test-time refinement. They report a 20.7-point gain over the strongest external baseline and release the framework, benchmark, and film KG dataset. ArXiv · AI/CL/LG's note
score 5