FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models

Lingjuan Lyu (Sony AI) · Hong Jia (University of Auckland) · Ting Dang (University of Melbourne) · Nicholas Lane (University of Cambridge) · Yan Gao (Chinese Academy of Sciences) · Massimo R. Scamarcia (ethicalabs.ai) · Javier Fernandez-Marques (Samsung AI) · Mohammad Naseri (Flower Labs) · Chong Ng (Flower Labs) · Dimitris Stripelis (Flower Labs) · Zexi Li (University of Cambridge, Zhejiang University) · Tao Shen (Zhejiang University) · Jiamu Bai (Penn State University) · Daoyuan Chen (Alibaba Group) · Zikai Zhang (University of Nevada, Reno) · Rui Hu (University of Nevada, Reno) · InSeo Song (Gachon University) · KangYoon Lee (Gachon University) · Junyan Wang (University of Adelaide) · Zheyuan Liu (Australian National University) · Daniel J. Beutel (University of Oxford)
aggregation strategiesbenchmarking suitecollaborative approachdecentralized fine-tuningdomain adaptationdomain-specialized llmsdomain-specific evaluation metricsfederated instruction-tuningfederated learningfine-tuning strategiesmodel performanceopen-sourcepre-trained modelsprivacy-preservingresource constraints

Large Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications.