By Dr Rohin Y, Founder-CEO, LightSpeed Photonics
Unlike the AI of ChatGPT & Gemini, we consumers are used to, enterprise AI requirements are quite different. Firstly, they have to be built to be reliable, and scalable, they should be maintainable for many years, to support their customers in a consistantly. It also has to have certain data privacy requirements, to securely store and process the data.
So, whatever planning that the enterprises are doing for creating an AI infrastructure, or any upgrade that the enterprises are doing for the AI infrastructure or there is a lot of emphasis on the data privacy, there is a lot of emphasis on how they will be able to scale it up and reliably serve their customers for long term to come. This way, Enterprise AI needs are very systematic and consistent.
And in fact, some of these applications are also used for their internal applications, repetitive tasks such as organisations are more and more creating agentic workflows that are not just acting as, systematic agents for defined tasks, but they’re actually able to perceive data across multiple platforms and create a multi-step projects and plan with very minimal human intervention.
As major enterprises move towards Agentic AI platforms and solve complex problems using layered AI sequences, there is more demand for edge “on-prem” AI hardware.
So that sort of an enterprise requirement, absolutely we need to have a good amount of compute happening on premise and there is a definite need to protect, for example, the coding that was developed for specific customers or data that is analysed for a certain organisation, all the code database that’s created through vibe coding thing or using specific small language models, built for high security, say analysing some legal contracts or sensitive data like that.
Basically, dealing with sensitive data and information, enterprises are more and more looking towards having as much, inference capability of these large language models on-premise, and continue to use the models developed at the hyperscalers, given to them by the likes of OpenAI, Gemini, and others.
So, with this, as a major trend, what we are seeing is the total cost of ownership for owning the hardware versus renting cloud hardware is becoming more and more lucrative to enterprises now.
And it is also becoming more meaningful in terms of owning and having more infrastructure where they can store data securely and protect the sovereignty and have the repository built over a long time giving them a specific edge in their niche areas of business.
From a purely development perspective, it may seem fine and even great! In fact, a lot of enterprises are excited to not have to pay for expensive NVIDIA hardware, like the B200s and 300s and the Vera Rubens.
As they can use more affordable language processing units, whether it is from a more open infrastructure, like hardware provided by AMD or the likes of models like Groq that can be run on a lot of low compute capable hardware.
And this compute, that is used for enterprise AI, has very much moved away from calculating the teraflops kind of infrastructure math to calculating tokens that the hardware can process per second per dollar. That is becoming a more valid metric compared to enterprise AI to what the compute capability that will be shown.
So better and better hardware is becoming available at a more affordable cost and probably not so power hungry also. So, a lot of enterprises are not necessarily worried about this, as the AI hardware caters their local needs.
From a naïve on-looker point of view, this arrangement looks fine, but there is a certain amount of discomfort from the hardware infrastructure point of view. As we see the architectures evolve, the long-term utilization plans and certain limitations become obvious that are very difficult to overcome.
And one very obvious thing, if you think a little deeper, as an enterprise accumulates a lot of data, it gets accumulated over the months, over the years, one starts realising where it can actually become a bigger bottleneck.
The software is very much coupled to the hardware with which the computation happens. And it’s not just the magic of how computing is happening in one place anymore.
Because you have a lot of this data stored in the SSDs and a lot of this data, you need to be able to move from, GPU to GPU to share between, if you need to do a compute that needs to have a very minimal latency.
So certainly, as the database becomes more and more massive, say with some historical context of logs that you need to be able to pull out or vector databases that are becoming more and more massive.
The retrieval of this data becomes a big bottleneck because the RAM cannot hold the whole data set all together for the compute to access readily. And no matter how much HBM you have on the GPU, you’ll still face this as a major bottleneck. And, this necessitates very low latency, high-speed connectivity from the storage SSDs like NVMes. So NVMe-over-fibre becomes a necessity. Now, it’s not good to have any more, it becomes a must have!
Similarly, if you think about some of these trends where people talk about in-memory computing becoming necessary because you will be able to, at least process some of this data near the storage or near the memory before it is sent to the computer engine.
So the data as it accumulates over many years starts becoming a big bottleneck naturally.
And if you think from further, you know, contextually, let’s say, let us look at a specific example for an enterprise use case where somebody has to sift through 1000s of documents and historical data, say the contracts that the company has done and need to generate something new.
And as years pass, these databases also become very large, like maybe some 10-20,000 PDFs to go through all in one go. So how many tokens can you realistically handle? And how much can you share between the memory and the compute?
Even when you have, HBM on-chip, moving it from the SSDs to cache, or memory like DDR 5, data movement becomes a very strong bottleneck.
And it’s basically like as if the, as the data volume becomes bigger and bigger, it becomes fatter and fatter. You can call it data gravity as it is difficult to push this data bandwidth to realistically do computations in real-time.
So there are better standards required for data interconnects and data transfer protocols. Enterprises are adopting low latency, high bandwidth connectivity such as the ultra-ethernet (consortium) to transfer data between the processors and upgrading Infiniband, which has been the gold standard in this domain.
And the hardware level, you know, the NVMe-over-Fiber alone is not enough, low latency memories are also required. So CXL, the Compute Express link, which was created as a consortium, is to address providing a huge amount of massive pools of memory, as external memory, as if it is sitting right on the compute engine’s, computer engine server, as if it’s its local data
And the GPU basically just needs to point to the right structure through this, you know, massive pools of CXL data. So this is necessarily a big challenge and if you go down further from the protocol level to hardware infrastructure in terms of the existing standards to further at the deeper hardware level, the data moves as bits and bytes through the SSDs that comes out of the chip through copper.
And that when it converts to optical signals to be able to send it fast across, there the integration of this with optics becomes an important bottleneck.
And the better that we are able to do this with low latency and adding high bandwidth, the more you will be able to cater and reduce this so-called interconnect bottleneck.
So effectively, you know, integration of high speed data transfer with photonics becomes necessary. It’s not just good to have anymore.
So electrical copper-based interconnects cannot definitely scale for providing connectivity between GPUs to memory, GPUs to storage or GPU to GPU, or even to enable some of these, disaggregated computing architectures that are required for the next generation computer systems.
All of these are going to become a bigger and bigger problem as the data accumulation happens on the inference engine hardware architectures, that it becomes a wall that is very difficult to break.
And the biggest issue isn’t the compute. Individually, you may be able to have teraflops of compute engines possibly available, but how quickly you are able to share data between these compute engines is where the future problems really lie, with the interconnect at its heart.
We, at LightSpeed Photonics, forecasted problems like these 5-6 years ago and created some of these solutions. Eventually, people developing the software get greedy, they want to do the work and vibe code what a 100-person team can do with just 2 or 3 people. They want to be able to generate software architectures from end-to-end. They want to be able to analyse huge codebases that are being created by AI further adding to the data volume.
So as the software gets greedier, the hardware pressure increases and the biggest bottleneck where this hits the wall and the interconnect problem becomes difficult to break.
Certainly, there is a lot of work happening on the networking and connectivity side. Both by the incumbents and upcoming startups, whether it is Silicon photonics, or alternative photonics, like what we do at light speed, photonics, and, you know, whether it is built as a generic copper-based system (that is definitely hitting a wall, both in terms of data transfer capabilities, the distance that can send the data normally). To the effect, CPO: Co-packaged optics efforts have started taking shape as optical engines, to integrate at the chip level. But it becomes very expensive and also very difficult to scale for, to create generic scalable, near packaged optical interconnects that have far wider interconnect bottleneck problems that it can address.
All these different approaches are going to create solutions for this interconnect bottleneck that is becoming increasingly prominent. And yes, the future is not all gloomy, as we also foresee that new computer architectures and novel hardware solutions will emerge from requirements like this in the next 5 to 10 years.
We will be able to create hardware that is beyond Silicon Photonics and standard computer architecture-based solutions to something that, the world of photonics can enable literally doing computations and communications at lightspeed to overcome the interconnect bottleneck for the future AI/HPC compute requirements.