Megadose AI progress, ranked and analyzed.

Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

· ArXiv · AI/CL/LG ·
Visual grounding helped the small model remember object properties, but standard BabyLM tests mostly missed it.

Lisa Bylinina seeds some DeBERTa token embeddings with image-region representations before training on 10M words. The visual initialization leaves a lasting trace, yet it does not move most abstract grammar benchmarks. The clear gain appears on object-property knowledge, and a tailored swap benchmark narrows that advantage to the seeded words. Synthetic grounding moves the effect to newly seeded words, while gains for function and abstract words lower mask-prediction loss without showing up in the tested benchmarks. ArXiv · AI/CL/LG's note

score 4

Categories: Research