Leaderboard · MultiFinBen

Open Financial LLM Leaderboard

Real financial work mixes languages and media: filings, scanned statements, earnings calls. MultiFinBen tests LLMs on 36 datasets in five languages and three modalities — text, vision and audio.

0models
0datasets
0languages
0modalities
Top 8 · modality-balanced scoreFull ranking →
Findings

What the numbers say

Ranking

Who leads, and in which modality

The modality-balanced score is the mean of a model's text, vision and audio averages, so a model that cannot see or hear scores 0 there. Switch to a single modality to compare only the models that support it.

Benchmark

What's inside MultiFinBen

Every tile is a dataset; its fill shows the best score any model reached. Hover for the top three, click to open it below.

Languages

Same model, different language

Average text score per language for every model that handles text. BI is the bilingual DOLFIN task; MU are the multilingual PolyFiQA tasks, which mix sources in several languages.

Datasets

Pick a dataset

Data

All results

Standardized scores (0–100) from MultiFinBen Table 3. A 0 can mean the model scored nothing or does not support that input.

People & partners

Built together

Author affiliations on the two papers behind this page.

Leaderboard · Open FinLLM Leaderboard: Towards Financial AI Readiness

Benchmark · MultiFinBen

Partners SecureFinAI Lab NaCTeM Archimedes AIRC

Method

Datasets are chosen with a difficulty-aware selection that keeps one dataset per modality, language and task tier. Evaluation code is in the MultiFinBen repository; scores are standardized to 0–100.

Add a model

Open an issue on MultiFinBen with the model's Hub id and revision. We run the full benchmark and add the scores here.

Cite

@misc{lin2025openfinllmleaderboardfinancial,
  title={Open FinLLM Leaderboard: Towards Financial AI Readiness},
  author={Shengyuan Colin Lin and Felix Tian and Keyi Wang and Xingjian Zhao and Jimin Huang and Qianqian Xie and Luca Borella and Matt White and Christina Dan Wang and Kairong Xiao and Xiao-Yang Liu Yanglet and Li Deng},
  year={2025}, eprint={2501.10963}, archivePrefix={arXiv}, primaryClass={cs.CE},
  url={https://arxiv.org/abs/2501.10963}
}

@misc{peng2025multifinbenbenchmarkinglargelanguage,
  title={MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application},
  author={Xueqing Peng and Lingfei Qian and Yan Wang and Ruoyu Xiang and Yueru He and Yang Ren and Mingyang Jiang and Vincent Jim Zhang and Yuqing Guo and Jeff Zhao and Huan He and Yi Han and Yun Feng and Yuechen Jiang and Yupeng Cao and Haohang Li and Yangyang Yu and Xiaoyu Wang and Penglei Gao and Shengyuan Lin and Keyi Wang and Shanshan Yang and Yilun Zhao and Zhiwei Liu and Peng Lu and Jerry Huang and Suyuchen Wang and Triantafillos Papadopoulos and Polydoros Giannouris and Efstathia Soufleri and Nuo Chen and Zhiyang Deng and Heming Fu and Yijia Zhao and Mingquan Lin and Meikang Qiu and Kaleb E Smith and Arman Cohan and Xiao-Yang Liu and Jimin Huang and Guojun Xiong and Alejandro Lopez-Lira and Xi Chen and Junichi Tsujii and Jian-Yun Nie and Sophia Ananiadou and Qianqian Xie},
  year={2025},
  eprint={2506.14028},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2506.14028},
}