Compute Infrastructures for IA applications in the wild - second edition

Workshop Details
Practical Information
None
Prerequisites: NONE
About this Workshop
With the advent of Chatbots, LLMs and other generative IA technologies, as well as other progresses in the IA field, there is an explosion of the demand for compute force. IA is no longer computer science: it is computational science. As such, it can no longer be done with casual, self-managed equipment. More advanced compute infrastructures are required both to satisfy user needs (in terms of compute power, GPU Ram capacity) and to ensure a decent utilization of the increasingly costly resources.
Content and topics
The purpose of this workshop is to gather people in charge of compute infrastructure (on-prem, cloud or hybrid) destined to support AI workloads (both training and inference). Being “in charge” means someone who does typically performs one or more of the tasks below
• Prepares and follows purchase orders for servers including GPUs.
• Racks the servers, ensure they are connected with adequate network / IO performance
• Install the OS, the drivers, some software stack
• Binds the servers with the login system, typically LDAP
• Monitor the utilization of the server over time
• Ensure there are no abuse
• Insert the server into a cluster (e.g. SLURM, K8s)
• Designs and implement cluster architecture
• Helps AI engineers to run their workloads according to the policies
• Provisions cloud GPUs resources
• Manages cloud platforms accesses
• Secures the cluster
• Dimensions the cluster and plan its evolution over the time
The workshop will be organized around 8-10 presentations, followed by a group discussion
The workshop has no specific registration, and walk-ins are welcome.
Schedule and speakers: to be announced by mid-February
Workshop Program
09:00–09:05
Sébastien Rumley, HES-SO
Introductory remarks
09:05–09:18
Vincent Magnin, HEIA-FR
Experiences with SLURM
09:18–09:31
Lucas Crijns, ArmaSuisse
SLURM with Open OnDemand for novice users
09:31–09:44
Pawel Bednarek, UNIFR
The DIT high-performance computing infrastructure
09:44–09:57
Alexandre Flament, HEG-GE
On-premises in a small research lab
09:57–10:25
Marco Merkel, HPC-DoItNow
Building the Foundation for Real-World AI From Ideas to Reliable Operations
10:25–10:40
Coffee break
10:40–11:05
Speaker to be confirmed, Exoscale
Title to be confirmed
11:05–11:18
Célien Donzé, HEIA-FR
Practical issues with ARM
11:18–11:31
Yann Sagon, UNIGE
GPUStack in production in our HPC cluster
11:31–11:44
Darko Petrovic, HEI-VS
OpenStack deployment with JUJU and MAAS
11:44–12:04
Franck Moreau and Anthony Glidic, Purestorage
Pure1 AIOps Never Sleeps (So You Can)
Speakers & Organizers
Sébastien Rumley
Associate Professor, HEIA-FR, HES-SO, Fribourg
HEIA-FR, HES-SO
Sébastien Rumley is professor of software engineering in the engineering school of Fribourg. Expert in large, distributed computing systems architecture, his research interest lay at the intersection of energy systems, IT systems, and sustainability.
Bertil Chapuis
Professor, HEIG-VD, HES-SO, Yverdon
HEIG-VD, HES-SO
Bertil Chapuis is a professor of software engineering at HEIG-VD. He helps teams design and build great software and systems in the age of artificial intelligence.