SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models
SLICEChat prunes whole-slide image tokens inside the encoder before multimodal fusion, cutting the expensive part of pathology VQA.
The paper says the model uses a hybrid Mamba-Transformer slide encoder to keep long-range context while progressively shortening the token sequence. Its pruning is language-supervised and region-aware, removing low-utility spatial regions on a controlled keep-rate schedule. On SlideBench VQA, it reports 79.84% accuracy on TCGA and 59.09% on BCNB, ahead of prior slide-level pathology MLLMs in the authors’ comparison. It also claims the highest overall WSI-Bench metrics, with competitive memory use and inference latency. ArXiv · AI/CL/LG's note
The paper says the model uses a hybrid Mamba-Transformer slide encoder to keep long-range context while progressively shortening the token sequence. Its pruning is language-supervised and region-aware, removing low-utility spatial regions on a controlled keep-rate schedule. On SlideBench VQA, it reports 79.84% accuracy on TCGA and 59.09% on BCNB, ahead of prior slide-level pathology MLLMs in the authors’ comparison. It also claims the highest overall WSI-Bench metrics, with competitive memory use and inference latency. ArXiv · AI/CL/LG's note
score 5