IndicBankBench: Evaluating Safety and Reliability of Language Model Assistants in Indian Retail Banking
IndicBankBench tests banking assistants where the failure can happen before the final answer.
The benchmark has 799 Indian retail banking cases across five operational domains, plus capability and refusal tests. It scores models at four stages: safety, tool use, response adequacy, and advisory quality. Across eleven models, strict reliability was 43.7% to 58.2%, while at-least-once success reached 60% to 74%, a gap the authors say can overstate dependable behavior. The authors release the cases, mock environment, and evaluation harness. HF Daily Papers' note
The benchmark has 799 Indian retail banking cases across five operational domains, plus capability and refusal tests. It scores models at four stages: safety, tool use, response adequacy, and advisory quality. Across eleven models, strict reliability was 43.7% to 58.2%, while at-least-once success reached 60% to 74%, a gap the authors say can overstate dependable behavior. The authors release the cases, mock environment, and evaluation harness. HF Daily Papers' note
score 4