The “Fraunhofer IIS Research Environment for Artificial Intelligence” (Fire-AI) is the dedicated GPU cluster that supports AI research within the DSgenAI project. Designed for the training of large language models and data-intensive workloads, it provides the compute, networking, and storage infrastructure required for modern AI research.
Compute, Networking and Storage at Scale
The newly built system consists of 93 compute nodes. Each node combines four high-memory GPUs with ARM-based CPUs, reflecting architectures commonly used in modern AI development environments.
For distributed training, communication between GPUs and nodes is often as important as raw compute performance. Fire-AI therefore uses a fully non-blocking interconnect with 400 Gbit/s per GPU, enabling high-bandwidth data exchange across the cluster.
The storage system provides 12 PB in total. Of these, 2 PB are NVMe storage for active training data and checkpoints, while 10 PB of disk-based storage are used for larger datasets and archives. The system reaches more than 1,200 GB/s read bandwidth and over 600 GB/s write bandwidth.
Fire-AI provides the infrastructure needed for working with large language models, especially for fine-tuning, adaptation, and evaluation. These workloads benefit from high GPU memory, fast interconnects and high-throughput storage. The cluster also supports multimodal experiments, preprocessing pipelines, and applied machine learning research.
Energy-Efficient AI Infrastructure
In addition to performance, energy efficiency played a central role in the system design.
The state‑of‑the‑art computing cluster uses warm-water cooling and redirects more than 95 % of the generated heat. Water is circulated through the system and an outdoor heat exchanger, eliminating the need for an active cooling unit and reducing cooling overhead compared with conventional air-cooled setups.
With Fire-AI, Fraunhofer IIS has established a research platform that provides the computational resources required for advanced AI research while also serving as a testbed for scalable and energy-efficient AI infrastructure.
