Viet Reader.

Loading...

Viet Reader.

VR.

Premier Newspaper for Vietnamese Worldwide

China Merchants Bank Wins CNCF End User Case Study Contest for Unifying AI Training and Inference on Kubernetes

China Merchants Bank Wins CNCF End User Case Study Contest for Unifying AI Training and Inference on Kubernetes

New cloud native platform lifted average accelerator compute utilization from 35% to more than 60% and cut inference cost per 1 million tokens by more than 60%

Key Highlights

  • China Merchants Bank, one of China's leading commercial banks, won the CNCF End User Case Study Contest for KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China 2026.
  • The award recognizes the bank's unified Kubernetes control plane built with Kueue, KEDA, Prometheus, HAMi, and Fluid that lets AI training, fine-tuning, and online inference share nearly 10,000 heterogeneous accelerator cards.
  • Bringing 99% of its accelerator compute resources under this framework lifted average utilization from 35% to more than 60% and cut the cost of processing 1 million tokens (input and output combined) by more than 60%.
  • The bank's in-house Twinkle training framework lets five LoRA tenants share one base model instance by default, cutting accelerator resource usage for that setup by 80% while increasing training density by fivefold.

SHANGHAI, Sept. 8, 2026 /PRNewswire/ -- The Cloud Native Computing Foundation® (CNCF®), which builds sustainable ecosystems for cloud native software, today announced China Merchants Bank, one of China's leading commercial banks, as the winner of the CNCF End User Case Study Contest for KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China 2026.


"As AI workloads continue to expand to critical infrastructure sectors such as financial, the core technology supporting it has to work harder without multiplying cost," said Chris Aniszczyk, CTO, CNCF. "China Merchants Bank's approach to unifying AI training and inference is a clear example of how composable, vendor neutral infrastructure powered by CNCF projects delivers measurable efficiency at scale."

China Merchants Bank's AI infrastructure team built a unified control plane by combining Kubernetes with Kueue, KEDA, Prometheus, HAMi, and Fluid, letting model training, fine-tuning, and online inference share nearly 10,000 heterogeneous accelerator cards. Through close collaboration with its data center and other teams, the bank brought 99% of its accelerator compute resources under this unified framework — lifting average utilization across those nearly 10,000 cards from 35% to more than 60%, while cutting the cost of processing 1 million tokens (input and output combined) by more than 60% under comparable model and service conditions.

"At China Merchants Bank, we unlock the potential of hardware through software innovation," said  PeiXiang Tan, AI Infrastructure Architect, China Merchants Bank. "Built on a unified cloud-native foundation, our platform brings together heterogeneous compute pooling, training job scheduling, elastic inference scaling, accelerated access to data and models, and end-to-end observability. This allows training and inference to follow workload-specific paths while working in close coordination—helping every unit of compute deliver greater value over time. "

With AI adoption expanding across more financial use cases, China Merchants Bank found that training, inference, and multi-tenant fine-tuning place significantly different demands on the same hardware. Distributed training requires stable, predictable capacity that can sit idle while waiting for remaining workers or accelerator cards to become available, while online inference needs to scale quickly with unpredictable traffic. To address these competing demands, the bank built a composable architecture that shares infrastructure while decoupling runtimes: Kueue manages training admission, queues, and quotas so jobs don't reserve capacity before they can use it; KEDA and Prometheus scale online inference from live demand signals; HAMi allocates shared accelerator capacity in fine-grained units across both paths; and Fluid accelerates access to datasets, model weights, and checkpoints so accelerators spend less time waiting on data.

China Merchants Bank's in-house Twinkle training framework adds another layer of efficiency for multi-tenant fine-tuning, letting five LoRA tenants share one base-model instance by default. Reducing base-model replicas from five to one cuts accelerator resource usage for this setup by 80% while increasing training density fivefold.

China Merchants Bank has published previous CNCF case studies on its AI infrastructure work, including one detailing a Kubernetes and HAMi-based scheduling platform that achieved 100% hardware pool utilization through topology-aware scheduling and a 30% reduction in cross-machine scheduling for distributed training workloads.

China Merchants Bank plans to continue building on these advances, focusing on four key areas: adjusting multi-tenant training concurrency dynamically, combining utilization, queue state, and latency signals into unit-cost-based capacity management, extending KEDA toward serverless inference that can scale to zero between demand spikes, and broadening HAMi-based support for heterogeneous accelerators while adding more training and inference backends.

China Merchants Bank's achievement was announced on stage during a keynote at KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China.

To review the complete architectural implementation and downstream results, read the full China Merchants Bank case study. 

Additional Resources

  • CNCF Newsletter
  • CNCF Twitter/X
  • Learn About CNCF Membership 
  • Learn About the CNCF End User Community

About Cloud Native Computing Foundation

Cloud native computing empowers organizations to build and run scalable applications with an open source software stack in public, private, and hybrid clouds. The Cloud Native Computing Foundation (CNCF) hosts critical components of the global technology infrastructure, including Kubernetes, Prometheus, and Envoy. CNCF brings together the industry's top developers, end users, and vendors and runs the largest open source developer conferences in the world. Supported by nearly 800 members, including the world's largest cloud computing and software companies, as well as over 200 innovative startups, CNCF is part of the nonprofit Linux Foundation. For more information, please visit www.cncf.io.

The Linux Foundation has registered trademarks and uses trademarks. For a list of trademarks of The Linux Foundation, please see our trademark usage page. Linux is a registered trademark of Linus Torvalds.

Media Contact:
The Linux Foundation
[email protected] 


Source: Cloud Native Computing Foundation

About author
You should write because you love the shape of stories and sentences and the creation of different words on a page.
View all posts