Infineon key visual in RGB

Health monitoring for reliable AI power distribution

We’re bringing you the latest from the world of semiconductors – straight to your ears! From quick takes on trending applications to deep dives on product innovations, our experts give you their take on the tech behind the tech.

Podcast

In this episode of Podcast4Engineers, we look at the role of predictive failure in the AI data center arena. How can we accurately predict which components will fail and when? Our conversation with Infineon System Architect, Christian Bernhardt will shed some new light on trimming the costs of ownership for hyperscale data centers.

In this episode of Podcast4Engineers, host Peter Balint speaks with Christian Bernhardt, System Architect at Infineon.

Peter Balint

Host:

Peter Balint has shaped visual and audio narratives at Infineon since 2021. He’s a video producer with 20 years of experience and has produced podcasts for the past 10 years. Over his career, Peter has interviewed speakers from all over Europe, bringing high-quality media production and engaging conversations to the forefront of his work.

Christian Bernhardt

Guest:

Christian Bernhardt holds a master’s degree in electrical engineering and began his career designing advanced measurement systems for automotive applications at IFTA GmbH. As a System Architect at Infineon, he brings extensive expertise in system engineering and embedded software development. Christian specializes in driving innovation in system failure detection and analytics, delivering cutting-edge solutions to meet complex industry challenges.

More episodes

Guest: The trend of AI really accelerated the importance of availability. One failure costs around €700,000 on average. You want to optimize the reliability margin of your whole system. We developed a solution called Power System Reliability Modeling. We really think like a stepwise approach is the key to implement health monitoring for customers.

 

Host: This is the Podcast4Engineers, the podcast you just have to listen to if you're interested in what's going on in the semiconductor market. I'm your host, Peter Balint, and we continue our journey into the world of powering AI. Today we're sitting with Christian Bernhardt. Welcome. He's a system architect at Infineon. Glad to have you here today.

 

Guest: Thanks for having me.

 

Host: You're welcome. You know, when we talked offline, you mentioned this concept of availability and the availability of servers and services. And I got to thinking, how does this feed into this whole idea of powering AI? Maybe you could bridge this for us.

 

Guest: Yeah, the trend of AI like ChatGPT and generative AI use cases really accelerated the importance of availability. And the reason is why is because the way how data centers are utilized changed with this new introduction. Like historically speaking, data centers usually hosted like websites and there were like web applications hosted on a lot of different servers around the globe. Like, and now if you think about Black Friday, for example, like the owner of those websites really wanted to stay them like that they are available now, but they could prepare for something like that. For example, as Black Friday was coming in, they bought extra bandwidth with a lot of different servers and therefore they were optimizing their reliability. And if the server went down, they had a lot of different servers as well where like the web application was still hosted. Now with AI use cases, you have basically for one individual AI use case, an orchestration of hundreds of different AI acceleration cards, and they are doing this in parallel. And if now one of them fails, if this computation fails, like 2 things could happen. Either the whole computation breaks down, and you have a faulty result, or you need to restart the computation. And those are very, very costly events.

 

Host: Okay. Now you talk about these failures. Does this happen often that there are failures?

 

Guest: Yeah, it's, you always have to put it into perspective. Like in general, failures in semiconductors and power supplies are very rare. But I think like Danny Clavette in a previous episode did a very great job in explaining how many power supplies you actually have in the field. Like just as a short summary, you have 800 data centers, soon to be 1,000, and in every data center you have like 10,000 racks. And within one of those racks, you have 10,000 power semiconductors and passive components. Now, one component of these fails could mean that the whole server goes down. And here it's because of the numbers, not that unlikely. And you have like, to give you a rough number at the Open Compute Summit, somebody told us that every 26 minutes, one power supply can fail. And one failure, which includes also a downtime, costs around €700,000 on average.

 

Host: Wow. Okay, so this is like the old school failures, but under magnifying glass.

 

Guest: Yes, absolutely.

 

Host: Incredible. What were some of the traditional ways of optimizing availability in the past?

 

Guest: Yeah, yeah. I mean, there are a lot of different angles you can tackle this issue, and a few of them Infineon doesn't even contribute to them. Like for example, like any complex machine, the operators need to be able to operate them. Human errors like cause a lot of failures. Therefore, like the operators are very well advised to train their staff to maybe not drink coffee right next to the AI server, right? But there are also like with the trends of AI servers where the power density is increasing significantly. This also has an implication that the temperature within these AI servers is increasing, and a general rule of thumb is that for every 10 degrees more temperature, the lifetime of your device will be halved. So thermal management in data centers is one of the key issues, even getting more important in AI servers.

Also, there are a lot of ways how to optimize reliability where we as Infineon play a very crucial point in enabling our customers. One, also, I can just recommend the podcast episode with Roberto Rizzolatti where he talked about the right choosing of power topologies, because the right power topologies have an impact on their efficiency, and the efficiency of the power supply also indicates how much heat is generated in the power supplies just because of inefficiencies essentially. So, and this excess heat needs to be got out of the power supply again. You're optimizing really by choosing your power supply reliability, by selecting the right topologies is key. But also, it's sound reliable in design and selecting the highest reliability components. And what is basically like an example of this would be that you want to optimize the reliability margin of your whole system. And the reliability margin is essentially like what you would expect your normal operation conditions and under which operating conditions your power supply will fail. The difference between those 2 operational points is your reliability margin. If you have a bus voltage of 48 V, you can choose a capacitor which is rated for 50 V, but I mean, you only have a reliability margin of 2 V. I wouldn't sleep well at night if I designed it like that. But what you can do is that you select a capacitor which has like 63 V, but this capacitor is more expensive, right? It's bigger. Therefore, it's always a trade-off with these traditional reliability methods basically between your cost and reliability.

 

Host: But as we have seen, the power supplies nowadays are getting much smarter. Is there any way to incorporate the smartness of a power supply into reliability measures?

 

Guest: Yeah, absolutely. I mean, digitalization enables new use cases across all industry, and this is also the same for reliability. Like 2 very, like almost buzzwords I would like to say, like are health monitoring and predictive maintenance. And both of those trends really can enable increase in the system's reliability.

 

Host: When we talk about this topic of reliability in an AI server environment, for example, we often hear this term health monitoring. Can you tell us a little bit more about that and what exactly is that in the semiconductor realm?

 

Guest: Yeah, absolutely. Health monitoring in itself is basically the system that tells you how healthy it is. Like a very like example where you can see this around is just in your smartphone. If you go into settings, you can see what the health status of your battery is. And this is basically the battery telling you, I'm that healthy. And now to enable this for power supplies, you need to basically have 2 very important aspects. One is you need to know how the power supply or the battery are operated, and the average over the operation is basically, we call it mission profile. You need to know the mission profile, and you need a very good description of what affects your reliability or your health of the power supply. We call that a model. So, health prediction model. And with those 2 aspects, you can make your health prediction. To give you like an example, like for the battery use case, for example, the operating conditions would be how often did you charge your phone? Like how deep was the discharge? What was the average ambient temperature and what was the peak current that you consumed? Like all of those factors affect the reliability of your battery or the lifetime of your battery. And now the reliability model, the health monitoring model basically describes how do you need to combine those individual parameters to get to the health status of the overall power supply.

 

Host: Okay. But how can you now use this health status parameter?

 

Guest: Yeah, yeah. Like the major use case is predictive maintenance. Predictive maintenance is essentially a data-driven method to optimize your reliability of the systems and optimize it for cost, essentially. Like a very like tangible example would be how you do your maintenance on your car. Like with a car, if you drive it and then it breaks while you are driving it, that's very expensive. You need to get the car to the workshop, it needs to be repaired, you need to have a new car essentially for the meantime. This is what we call reactive maintenance. Then we have periodic maintenance, and periodic maintenance is essentially once a year you go to your, to your car dealer and do a checkup. And so here you basically improve the reliability of your car, but it might be actually not necessary. And then we have prescriptive maintenance. Prescriptive maintenance is the car manufacturer telling you every 120,000 km you need to change your brakes, and then you do that. Predictive maintenance now would be basically a sensor over your brake telling you exactly when your brake in a state is that you need to do your maintenance. This could be either 140,000 and you save on maintenance costs essentially, or it could be like 100, and then you need to go to maintenance early to prevent actually a critical failure.

Same logic applies to power supplies. And another case that we see, which becomes more and more important over time, is that a health status enables circular economy. Like, because if now the power, like the power supply vendor or any, like any OEM gets his system, he needs to make the decision, do I want to resell this? Do I want to need to repair this first, or do I just recycle it? And therefore, he needs to know how healthy the system is actually, what makes the most sense for me. And so here using this health status parameter can enable circular economy use cases. And in the end, all of them play into optimizing the total cost of ownership. Because with very effective health monitoring and effective maintenance strategy optimization, you can reduce your capital expenses as you need to buy less equipment, and you optimize your operational expenses because you prevent failures from happening and you optimize your maintenance.

 

Host: Okay. I'm wondering technically speaking, what does this health monitoring solution look like actually?

 

Guest: Yeah, yeah. The devil really lies in detail, and it can get very nuanced, but generally speaking, you have 2 different types of maintenance or the first important distinction that we have to make. One is you have it locally or the other one is you process it centrally. The difference essentially is that on a local system, like for example, a wind turbine, the wind turbine is checking itself, like what is my health status, for example? And then when it sees, hey, my health status is not very good, it might reduce the maximum power output to ensure I don't break until the next maintenance interval. And then the decision is being made like, okay, I do maintenance now. The central way would be that the wind turbine is quote unquote stupid and it's just broadcasting their data to like a management system, to like a central, like to a cloud. And then on the cloud there, the operator does the calculation to detect health status. It's basically like on the edge and the other is in the cloud, right? And those would be like the 2 major differences. For the local processing of health status, we call it in situ because it's happening in the device. And ex situ is basically then the reliability model or the health prediction is happening outside of the device. So therefore, ex situ.

 

Host: Okay, so you've explained ex situ and in situ, but what are the implications of each one?

 

Guest: Yeah. Essentially in situ here, especially in the realm of power supplies, you have the major benefit that you have all your operational profiles in very high resolution. Like, you know, all your system parameters. The drawback to that is essentially that you don't have that much computing power, so you cannot run the most sophisticated and complex algorithm to actually process this data. But on the other hand, like ex situ methods, you have all the computing capabilities that you require, but what you don't have necessarily is very high-resolution data. It's like, not one has like the, by far the upper hand. It's just like, how do you want to do your health monitoring? And like the 3 major methods how you do your health monitoring would be essentially statistical, would be based on physics on failure, and would be anomaly detection. And those, like maybe I will briefly go through all the 3 methods. The statistical model is that you collect a lot of data and then based on like mission profile data, failure data, and then you analyze this data, you know, based on the normal usage, when do my failures happen? And then you make a statistical analysis out of this. And then like you do your health prediction that you essentially say, I measure now the behavior of my current device, or I measure how my device has been used. And then I compare it to my expected statistical behavior. And based on this comparison, you can make an estimation how healthy it is. Physics of failure here is a method where you really need to know everything about your component. It's essentially like, I know my power supply breaks exactly after this time because like the solder cracks and you know, okay, my solder cracks after a certain number of thermal cycles. And now in your device, you count the number of thermal cycles and then you, in your lifetime model, you basically say, I had this number of cycles, now my device is this healthy. Actually, but getting to this level of understanding about your system, it's really complex. Last one is anomaly detection. And anomaly detection essentially says, this is now the baseline behavior that I would expect from my device. And now something has happened and now the device is working abnormally. Like it's not working the way I think it should behave. And then by the comparison of those 2 behaviors, you can say whether your device is healthy or it's not healthy. Thing is here, also very complex to determine what is normal behavior and what is abnormal behavior. That's the devil that really lies in detail.

 

Host: Yeah. Okay. Several approaches, but can you give me the global view now of what Infineon can offer in this area of maintenance?

 

Guest: Absolutely. Absolutely. So, we developed a solution called Power System Reliability Modeling. What we do there is a statistical approach which is in situ executed. And the way how our solution works is because we have the fit parameter of all the components that you use in your design. And every fit— so fit parameter is failure in time. And every parameter now ages. faster or slower depending on how it's used. Our digital power controller, which is basically the product of power system reliability modeling, is measuring in real time what the temperature, what the electrical stress, like the voltage and the current is, and accelerates each and every single component based on those measurements. And then we can calculate the overall failure rate of the power supply. And this power supply failure rate parameter now can be used to do a health estimation for the whole power supply. And for us, it was really important that we set a focus here on availability, essentially like that it's easy to implement. So therefore, our solution is topology and power level agnostic.

 

Host: Okay. So just another example of how Infineon is committed to powering AI.

 

Guest: Yes, absolutely.

 

Host: Okay, good. So final words from you.

 

Guest: Yeah. What I think, like health monitoring, especially in the beginning, is a very complex application, which has huge upside when it's implemented correctly. But this initial level of complexity is overwhelming a lot of customers, a lot of people. We really think like a stepwise approach is the key to implement health monitoring for customers. From 0 to 100, it will most likely be a bit too complicated, but with our solution, you get like the very first, very good health monitoring solution, which also helps you to understand your data, your operational of your power supplies even better. If you are interested in the topic, would be very happy to support you on this journey to really enable predictive maintenance for your systems. And if you're interested in it, like in general into a deep dive, feel free to visit our website where we have a lot of collaterals available where you can go in more detail.

 

Host: Excellent. I say thank you for coming in today and sharing your knowledge with us and with our audience. And I also say thank you to our audience for listening to today's program. We're constantly looking for ways to bring you better content. If you have any ideas, please send them to wepowerai@infineon.com. Thanks again, and we'll see you soon.