Hugo Scott-Gall: Hello and welcome back to The Active Share. A few weeks ago, I sat down with Jay Kannan who made the case that the most interesting investments in tech right now aren’t at the AI headline level, they’re actually one or two layers below it in the physical infrastructure that makes all of this possible. If you haven’t listened to that episode, I would recommend going back to it because today we’re going deeper.
Joining me now, joining me today, is Siuchoon Koay, a research analyst on our team who covers large-cap global semiconductors. Siuchoon came up through the engineering side. He has a degree in electrical engineering and a degree in computer science and economics from MIT. So that’s two degrees, which means he’s one of those rare people who can read the physics of a chip and the economics of a supply chain at the same time. Jay gave us the map, Siuchoon is going to take us inside it. Siuchoon, welcome to the show.
Siuchoon Koay: Hey, thanks, Hugo. Delighted to be here.
Hugo Scott-Gall: Excellent, right, let’s get going and talk about probably…roughly about the most interesting thing in the world right now, which is there’s huge demand for chips and the supply side, the manufacturers of chips, can’t make enough of them. So, I’m going to ask you this question, which is really about bottlenecks. Why can’t the makers of chips make enough to meet demand?
Siuchoon Koay: Yeah, I think, simplistically, I think it’s just how sudden all of it happened. I think, in terms of timeline, I think AI started since twenty, call it twenty-three. But really, I think demand took off quite recently, in twenty-five. And that’s where the addition of supply has been a lot slower, with the chipmakers, especially the foundries like TSMC. You know, especially after inferencing started.
So, I think that’s where there is a mismatch out there between supply and demand. Just that it suddenly took off and there’s just not enough time to pull together supply. And it’s not just actually with foundries, but also with memory, with substrates, with a lot of the other components that go into the server rack. So, I think that’s gradually being resolved as we speak, but it’s still not quick enough for that demand profile.
Hugo Scott-Gall: You’ve been doing this a long time, so you’ve seen a lot before. Have you seen anything like this, where things are so tight in the supply chain?
Siuchoon Koay: I’ve never. So I think in the past, you know, it was more a consumer adoption cycle with internet and smartphones. And this time around I think it’s just, the consumer is not the direct driver of that demand profile. I think it’s more AI frontier models and the cloud. I think those two pieces can run a lot faster than consumer adoption. So, you know, the simplistic answer to your question is no, we haven’t seen this sort of demand profile before.
Hugo Scott-Gall: If my chronology, my history timeline, is off here, but October of 2022, ChatGPT becomes the catalyst, the Cambrian explosion, the moment when the use case, I think, for GPUs to be useful in AI, that AI in its current form arrives. Is that moment where this all goes nuts in terms of demand and then stresses the supply chain?
Siuchoon Koay: Yeah, I think so. I think we’ve never seen this sort of tipping point appear before in other applications. So I think it was more gradual. And this time around, to your point, that explosion, that tipping point, did suddenly get turned on, right? From an off to an on switch with Gen AI.
So I think that’s where demand just took off, primarily in AI chatbots. And now we’re seeing coding as the other major example of usage. But I think going forward that continues, we will see, I think, more of graphics, right? Just AI generating photos and potentially AI generating videos as well. So I think that’s where it gets a lot more interesting, beyond text, in the first phase.
Hugo Scott-Gall: And if you could maybe sequence for me the phases we’ve had when it comes to where the bottlenecks have really been. Am I right in saying that the first bottleneck was GPUs, and demand for them was off the charts, and that led to Nvidia, several times, quarter upon quarter, having very, very strong financial results, very strong earnings that were well ahead of expectations, because their pricing power was so strong, because demand was so strong? Where are we now? Where have we moved on to? So if it started with the GPUs, the graphics processing units, where are the real, super-tight bottlenecks now?
Siuchoon Koay: Yeah, maybe starting with GPUs. I think that started the AI revolution with training, right? So I think that was the first phase, that takeoff, with training and GPUs. So as that extended itself into the second and the third year, we started to see more inferencing emerge. And inferencing is where usage comes from, right? Driving more sort of consumer and enterprise usage. So inferencing, I think, is several magnitudes bigger than training, because that is where it is really deployed out into the market. And as inferencing grows, we would require a lot more chips, not just GPUs, but also other form factors such as CPUs, LPUs, ASICs, and also everything that accompanies those chips, which is memory and substrate and the racks.
So people were quite unprepared for the demand profile, which we mentioned starting in 2025. And that’s when, you know, the supply adds have been a lot slower. And today I think that’s most acute in memory. We see that most acute in HBM. Just because with every unit of GPU that is deployed, you need also corresponding HBM, and that supply for memory was just not ready to take off in the same fashion.
Memory, I think, is also taking up more of the total capacity, the pre-existing capacity that’s available on the DRAM side of things. And I think that has driven up pricing a lot. So we’ve seen commodity DRAM double, if not triple, in pricing in 2026, first half. And that in turn has looped back to drive even higher prices in HBM itself.
So the shortage has kind of had a boomerang effect, right? So that bottleneck is getting resolved, but just not quite fast enough.
Hugo Scott-Gall: So we could spend time on each of the different parts of the supply chain. You talked there about memory, and there are lots of individual bits with high specialization from companies that dominate those niches. But I think maybe more interesting is just to go up a level and say, okay, how do you solve these bottlenecks? So you said memory is very tight, and you mentioned high bandwidth memory, HBM. So how do these things get solved? Certainly when I started my career, your industry, semiconductors, was always seen as not a great industry—very cyclical, and the cyclicality of demand was amplified by when companies added capacity and usually got it wrong. And so you had falling demand with growing supply, which meant prices collapsed, profitability collapsed. And this wasn’t really a great place to invest versus other sectors. That’s changed.
The biggest thing in the equity markets for the last 12 months has been the performance of the semiconductor industry, of the chipmakers. How do these supply bottlenecks get resolved? If I could help frame your answer—is it that the companies do it themselves, which is they just start buying more equipment so they can make more, or is it that there’s innovation, and you start seeing substitute products or alternative products created to get around these bottlenecks?
I want to ask you after this a bit more about company strategy, but if we accept we’ve got very, very tight supply, how does it resolve? Remembering, of course, that maybe it doesn’t, if demand just keeps getting stronger and stronger and stronger, supply still can’t keep up. But we can get onto that in a minute—how do the companies resolve this, or does technology innovation do it for them?
Siuchoon Koay: I think that’s a great question. So I think if you look back six months ago, some companies were not convinced about AI, right? So they were dragging and delaying whether or not to add. So I think that was one part of that bottleneck equation—that companies were not convinced. They were asking the question, are we in a cycle? Are we in the late stage of the cycle? Should we be adding that much in capacity, because it would create a glut down the road? So as the demand strength continued over the past six months, up until the present moment, I think companies, foundries and memory vendors, became more convinced that the strength would last longer. And, of course, by the time they realized that, it was a bit too late, and we’re just in a very tight situation as a result.
But to your second question—how does the industry resolve itself? I think everyone wants to bring costs down today, right? Not just with memory, but also with other parts, including the chip, including just the equipment required to make these things happen. So everyone’s looking for new design methods and solutions to bring costs down.
And as a result, we may see a proliferation of form factors, not just the really expensive ones, but also maybe second-tier, cheaper kits. ASICs would be one of them. So I think we see that you don’t need to always use GPUs. For some workflows you can be resorting to ASICs, and for others you can use CPUs, which are a lot cheaper than a GPU.
And the same thing applies for memory as well. Not always HBM, but sometimes you go down to either DRAM, or SOCAMM, or something of that nature, to bring the cost down eventually for the entire hardware stack. So I think this is how the industry today is trying to get at that same problem, which is that supply is extremely short. It’s very expensive, and it’s not going to make economic sense if we use the most expensive kit to do everything.
Hugo Scott-Gall: But that can only do so much. Profitability—you look at margins, or the absolute amount of dollars of profit—it’s never been this good in most parts of the supply chain. So surely, yes, they will try to do the same thing cheaper, they’ll innovate themselves, but aren’t they just going to have to spend a lot more on capex? They’re going to have to buy, wherever they can, more production capacity to make more stuff. And is that then the beginning of the end?
Siuchoon Koay: So I think the demand profile doesn’t seem to be slowing at this moment when we look at token demand specifically. It’s essentially growing at, call it, eight times to ten times every year, right? So that demand profile is very, very, very strong.
Now, I think as we go along, no one’s able to spend eight to ten times more capex every single year just to fulfill that. So there needs to be innovation, right? At the hardware level, at the software level, the algorithms—to bring some of these prices down. So every iteration we see today, in, for instance, Nvidia’s kit, hardware is trying to solve for lower cost, lower cost of compute, to bring token cost down. And that’s one way to get at the problem—that we can have more tokens at a cheaper price. Ultimately, it’s to make it cheaper so that we can have more adoption and see proliferation.
Hugo Scott-Gall: So, you’re not too worried that ultimately the cure for high prices—which is a problem for those who are buying the chips, since they’re getting more expensive—is a lot of capex to increase supply, which means this is going to be like every chip cycle, where in the end the makers just expand capacity too much? You don’t think that’s an imminent threat or risk?
Siuchoon Koay: No, I think what ultimately breaks is less from the demand side. I think everyone wants just, you know, infinite amounts of tokens. I think that’s very clear—that the AI use case is relevant and will continue to be relevant.
I think what breaks is potentially who’s paying for this. It’s more from that angle—the supply of financing, the supply of chips, the cost of the chips. I think that’s where the break ultimately happens for this cycle, if we may call it that. And today I think a lot of it, so far, has been coming from private financing. Some of it has been circular, as we’ve seen last year in 2025. IPOs helped a bit, in releasing this valve of requiring continuous NVIDIA funding from customers.
I think ultimately, as these frontier model companies go IPO, the market will be the one that supports the entire AI financing side of things. But longer term, we still need to answer the question of what margins are going to look like for frontier models.
And what is the disruption TAM—the total addressable market that AI can disrupt? I think these two questions would dictate how much valuation would be ascribed to the frontier models, and ultimately whether they could continue funding that chip supply, because these companies so far have been relying on private financing, but they’re going on to IPO.
Hugo Scott-Gall: Yes, and so that gets us to monetization—you buy a lot of chips, can you make money out of them? And that certainly seems to be a bigger question mark when you look at consumer than when you look at enterprise. When you think about how demand might evolve from agentic AI—lots of mini AI agents doing lots of work for people, whether that’s for companies, for individuals, for governments, or whoever—what does that mean in terms of the type of chips that are needed? If the average person has five, six, seven different agents doing things for them, organizing their lives, doing projects, doing presentations, and whatever their job involves…
…what sort of chips are needed for that world? That world feels not just imminent, it feels here—but of course here today may seem very, very basic compared to where it will be in three years or five years’ time. What does that mean for the type of demand that you see coming?
Siuchoon Koay: Yeah, I think it’ll be a lot more diverse than in day one. For agentic AI we anticipate that there will be a mix of GPUs, ASICs, and then a bit of LPU, and also a bit of CPU. So I think it’ll be a very mixed and diverse picture in terms of chip demand.
I don’t think it will be as pure as in day one, when a lot of it was training that relied primarily on GPUs. So I think the picture gets a bit more competitive for the players involved—everyone’s trying to outdo each other. CPU vendors are trying to innovate to also sell GPUs. IP makers are also trying to sell CPUs. And GPU players are diversifying into LPUs, and so on. So I think this picture gets a bit more competitive in day two, as we move into inferencing and agentic.
And that’s where it’s not so clear what sort of terminal value could be ascribed to each player in year ten, because I think winners and losers are still yet to be made out very clearly today.
Hugo Scott-Gall: There are some things I guess that have to happen at the moment in the world. There’s really only one dominant foundry for making chips—Taiwan Semiconductor Company, TSMC. You just outlined all the different types of chips—ASICs or custom chips, GPUs, CPUs. But there are some overlapping areas of opportunity, in that there’s still a dominant maker of chips and still some niche dominators in the supply chain.
I want to pivot now and talk about China, because I don’t know if most people have heard of DeepSeek—it was a kind of overnight sensation. The Chinese AI ecosystem seems happier to be a fast follower of the leading frontier models, the LLMs—so that’s OpenAI, that’s Anthropic, and a few more, Gemini from Google.
China seems happy to be a fast follower and have much lower costs of development and operation in its AI models. Is that a medium-term threat? Is that a better way of doing it—getting the majority of the capability for a lot lower cost? And what does that mean for the supply chain? What does it mean when China is seeking to have its own supply chain, with no real external dependencies? So the Chinese ecosystem may well look quite different and may well be quite self-sufficient. Is that a threat to the overall supply chain, to a lot of the companies we look at today?
Siuchoon Koay: Yeah, I think that’s a great question. China has been sort of pushing the open-source LLM model. A lot of it has been initiatives from the big internet companies, together with DeepSeek. So I think that’s the direction of travel today from the Chinese AI efforts. And they are a lot lower cost, to your point, because they’re using essentially older chips, and energy is a bit more abundant. So the total solution comes up to a lower cost per token on the Chinese side. Longer term, I believe that in critical applications in the US and Europe, a lot of that will still rely on frontier models coming out of the West.
I don’t think that will change imminently. Using open source still has issues in terms of trust, and how good these models are for mission-critical tasks. But longer term, as we pivot more to the edge—meaning every single household item could have intelligence—we could see more adoption of open-source models. And it may not be Chinese models; it could be, say, Nvidia’s open-source LLM. So I think open source may become more prevalent at the edge, in non-mission-critical applications.
That could challenge the pre-existing narrative that frontier models can make very rich margins all the time, because open source would present a margin challenge to them. Longer term, outside of the West, emerging markets could see higher adoption of Chinese models, which could pivot demand away from the names we traffic in today—the Western, Japanese, Korean, and Taiwanese AI semiconductor names—and favor the Chinese hardware ecosystem instead.
Hugo Scott-Gall: So the solution to today’s supply chain bottlenecks could either be the companies involved—mostly in Taiwan, Korea, and Japan, a bit in the US, a little in Europe—spending a lot more on capex to increase supply, or innovation to do the same thing better, cheaper, more efficiently. But then there’s also China as a medium-term threat. How much should people think about that? If you look at a lot of industrial end markets, China, through scale and innovation, has become really, really good, and has brought prices down, created capacity, and in many ways dominates the profit pool, even though that pool is probably smaller because China is operating at lower costs.
Is that the end game, do you think, for a lot of the semiconductor supply chain that’s outside of China currently—that China will end up the dominant, at-scale, lowest-cost producer?
Siuchoon Koay: Wow, that’s a really difficult one. If the Western frontier models cannot defend margins in the face of open-source models, it will become very untenable, because who’s paying for it? That question cannot be answered.
So, longer term, for hardware investors, you’re right that commoditization is the biggest fear, because things do get cheaper and cheaper to produce at the margin over time, over a decade or two, as we’ve seen. And that would erode the competitive advantage that players hold today. So, ultimately, the question boils down to whether frontier models can hold the attention of their user base in the West and defend margins. Because if that model falters, and open source becomes a lot more prevalent, then that competitive line could shift further into China’s favor.
Hugo Scott-Gall: So when you’re thinking through the risks to demand that could upset the very favorable demand-supply balance for the supply chain, the big thing you’re looking at is really the revenue-generating opportunity for these frontier models, the best LLMs—OpenAI, ChatGPT, Anthropic, which is more business- and consumer-focused, Gemini, and a few more. To justify using the best chips to have the best models, you’ve got to have revenue generation.
Have I summarized rightly—is that the biggest proof point here, because if the revenue generation is there, then everything else falls into place? The hyperscalers, the Googles, the Microsofts, who are spending a lot on building out cloud—that’s fine, because their customers can pay for those cloud services, so their capex makes sense. Is that where you see the central risk to the whole edifice—the monetization ability of the frontier LLMs?
Siuchoon Koay: Yeah, I do. I think it’s an important element for the AI stack economics to work out—that frontier models need to maintain their margins, or at least have an improving margin profile from here, and not see a declining one. Because if they can’t defend economics, even with volume growth, it would become very challenging down the road, because their valuations would come down, and every incremental dollar, or at least cash flow, that flows through would become more constrained.
I think NVIDIA is trying to hedge against this by offering its own LLM as an alternative to these really expensive frontier models and also to Chinese open-source models. So it sees this potential problem that people don’t necessarily want to use Chinese open source, so it’s offering this alternative, a US-based open-source model. So, ultimately, we need to see margins being defended at the frontier level. If not, I think investors will potentially question what terminal value will look like.
Hugo Scott-Gall: Okay, a couple more questions. In a minute, quantum computing—but just before that, let’s do corporate strategy. If you’re a big purchaser of chips, or a maker of chips, or you own the IP to develop chips, and you’re faced with these rising supply chain costs… you mentioned earlier some concerns around circularities of vendor financing, meaning someone like Nvidia investing into either its customers or potentially its suppliers. Is that a strategic choice companies today could consider—that an Nvidia, for example, could become tighter with its suppliers by investing into them, or JVing around R&D? That’s the case for vertical integration. Because this cycle is such a super cycle, such a long cycle, is that something you think can happen?
If I could add on to a question that’s already too long—there’s a prisoner’s dilemma happening here too. If you’re a supplier where demand is exceeding what you can supply, your choices are to increase your own capex so you can increase your own supply capacity, or stay disciplined but watch someone else do it for you. So it’s interesting to think about the strategic choices sitting with CEOs like Jensen at Nvidia—thinking about his customers and his suppliers—and then the strategic choices of, say, a TSMC, the world’s biggest maker of chips.
How disciplined do they stay on capacity, versus the risk of demand overflowing to others, which creates profits for those others and gives them the cash flow to reinvest? There are some quite tricky decisions going on in this super cycle. Is that fair?
Siuchoon Koay: Yeah, absolutely. I think investors have asked the question of whether TSMC has missed the boat by not adding capex soon enough. And a lot of investors would agree with that, which is why I think Intel and Samsung have been rallying—they’re seeing new customers sign on, and those new customers will improve utilization and, in turn, potentially create a virtuous cycle for the second- and third-place players in foundry. So by delaying capex adds, it does create a problem for the incumbent, because that supply bottleneck favors the number two and number three players as a result.
To that earlier question of whether NVIDIA ought to be investing in its supply chain—I think the answer is yes, because it is very tight today. By offering a financing option for capex adds to suppliers, Nvidia can essentially resolve some of these bottlenecks, and that gives it potential competitive time to market, which could give it an edge relative to competition. So, in a nutshell, I think Nvidia’s strategy is correct—it ought to be extending more supplier financing, and incumbent players also need to be more aggressive with their capex adds to maintain their competitive position.
Hugo Scott-Gall: Okay, final question on the risks. Our starting assertion was that we’re in a super cycle, the likes of which you—with more than 20 years of experience—haven’t seen before. So what’s happening now carries big capital lessons. But quantum computing has to be a big risk here. Rather than asking when quantum computing is arriving, what are you looking at that would increase the probability of quantum computing as a risk? What are the signposts? What is it you need to hear? And correct my assertion if I’m wrong.
Quantum computing is a big deal in that it changes how the digital infrastructure will work, and the nature of compute power itself. So what are you looking for? What’s your understanding, and am I right to be worried about this?
Siuchoon Koay: So we’re still looking for use cases for quantum computing. I think today, Jensen has mentioned before that QPUs would exist alongside GPUs and CPUs…
Hugo Scott-Gall: Sorry, what are they?
Siuchoon Koay: Quantum Processing Units. So these quantum computers would exist along GPUs and CPUs, because they’re meant for different types of workloads, right? What a quantum computer can do is not quite universally applicable to what parallel computing and serial computing, in the case of GPU and CPU, would do. So there’s still a bit of differentiation between use cases, at least from a quantum computing perspective.
So from that angle, the threat to GPUs and CPUs is maybe not so imminent, but I think that continues to be an evolving landscape, as algorithms improve for quantum computers and use cases broaden. Today it’s still very narrow, I’d say, the type of problems being solved by quantum computers, but as that broadens, it will become potentially more of a threat to classical computing. We’re not at that cusp yet—where you’d say GPUs and CPUs are over, that it’s the era of quantum computing. But going forward, that tipping point could happen, and we’re watching for it.
Hugo Scott-Gall: What are you watching? This all kind of happens in secret.
Siuchoon Koay: Well, there are some smaller companies that are already public today. So they do have sales, and they talk about who their customers are and what they’re buying these systems for. So we’re tracking that narrative to see whether it encroaches upon classical computing—whether these quantum systems are doing exactly the same tasks that classical computers can do.
So far it hasn’t come to the point where it’s completely replaceable. But people conjecture it could happen in the next five years. That continues to be a work in progress in terms of what quantum computing can do. So we are watching that space.
Hugo Scott-Gall: As I said several times, you’ve been doing this a long time, and you have high levels of expertise. This must be the most exciting time ever to be a semiconductor analyst.
I think it's super exciting. I think you’re in a great position to cover the picks and shovels, the enablers of this step change in the capability of compute power. Who knows what the next one, three, or five years looks like, but it’s super exciting and exhilarating.
So, Siuchoon, thank you for coming on the show, thank you for sharing some of your expertise, thank you for not using too many acronyms. But thank you, most of all, for setting out your framework for how you think about what’s happening.
Siuchoon Koay: Hey, thanks for having me here, Hugo.
This content is for informational and educational purposes only, and is not intended as investment advice or a recommendation to buy or sell any security or to adopt any investment strategy. Investment advice and recommendations can be provided only after careful consideration of investors' objectives, guidelines, and restrictions. The views and opinions expressed are those of the speakers as of the date of this recording, are subject to change without notice as economic and market conditions dictate, and may not reflect the views and opinions of other investment teams within William Blair Investment Management.
Factual information has been obtained from sources we believe to be reliable, but its accuracy, completeness, or interpretation cannot be guaranteed. Any discussion of particular topics is not meant to be comprehensive and may be subject to change. This material may include forecasts, estimates, outlooks, projections, and other forward-looking statements. Due to a variety of factors, actual events may differ significantly from those presented.
Past performance is not indicative of future results. Investing involves risk, including the possible loss of principal. Any investment or strategy mentioned herein may not be suitable for every investor. References to specific companies are for illustrative purposes only and should not be construed as investment advice or a recommendation to buy or sell any security. William Blair Investment Management may or may not own any securities of the companies referenced. It should not be assumed that any investment in the companies referenced was or will be profitable.