How AI could help data centers do more with less power

MIT researcher Christina Delimitrou applies machine learning to improve how data centers use servers, networks and shared computing resources. Her work also uses AI to anticipate application problems, helping reduce downtime and make existing hardware go further.

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The story describes AI automating data-center management, with a mild lean toward greater system autonomy and control.

How AI could help data centers do more with less power

Data centers are expanding to meet rising demand, putting pressure on electrical grids and increasing reliance on polluting fossil fuels. Christina Delimitrou, a newly tenured associate professor at MIT, is working to reduce that strain by changing how the computers inside these facilities use their resources.

Making existing data center hardware work harder

Delimitrou and her research group use machine learning to improve the efficiency, security and reliability of large-scale data centers. Their work includes updating cloud computing systems, managing shared hardware resources and designing more streamlined server architectures.

The goal is to get more computing capacity out of equipment that is already in place. If servers and other resources are poorly utilized, operators may use more power than needed to meet demand, and could face pressure to build additional facilities.

Delimitrou describes software bloat as one source of inefficiency. Removing unnecessary complexity without harming performance could let data centers serve users with fewer additional buildings. That makes software design and resource management relevant to the physical footprint and energy demands of computing.

Using machine learning to manage complexity

Managing a large cloud system involves many moving parts. Machine learning can automate some resource decisions and identify options that developers might overlook, particularly when the system is too large for people to tune every part by hand.

In earlier research with her mentor, Christos Kozyrakis, Delimitrou found that many large computing systems were using only about 15 percent of their capacity. That gap between available resources and actual use pointed to an opportunity: improve utilization before adding more hardware.

Her group developed Seer, a tool that uses deep learning to predict and prevent problems in web applications. Catching trouble early can avert broad slowdowns that might otherwise follow a manual fix, while avoiding downtime that wastes computing resources.

These systems also affect the experience of people using cloud services. More effective resource management can help applications on smartphones deliver more predictable performance, including services such as music streaming and video conferencing.

Adapting cloud systems as applications change

Cloud applications have shifted toward designs that split software into smaller components distributed across multiple servers. This can speed up deployment, but the servers were not built for this newer way of organizing applications.

Delimitrou adapted her research to address that mismatch, developing machine-learning approaches for the newer class of applications. She also works on redesigning software to better fit the capabilities of existing hardware, connecting software choices directly to how efficiently equipment can be used.

Her work has broadened from performance and resource use to security. She extended application-debugging research to cover security problems that could leave user data vulnerable to hackers. In this view, making cloud systems better involves identifying both operational faults and weaknesses that put users at risk.

A path from computer engineering to cloud research

Delimitrou grew up in northern Greece, where she became interested in mathematics and science. Her mother was a chemical engineer and her father a pharmacist, and both encouraged her curiosity. She studied computer engineering at the National Technical University of Athens, where a final-year project on managing resources across several applications introduced her to challenges that become more difficult at large scale.

She later pursued graduate study at Stanford University and began researching inefficiencies in cloud computing and data centers. After earning her PhD, she continued the work as an assistant professor at Cornell University, then joined MIT as an assistant professor in EECS in 2022.

At MIT, the research combines hardware and software perspectives. The central question remains practical: how can computing systems meet growing demand while using the servers and energy they already have more effectively?