ZML, a French AI startup, has launched ZML/LLMD, software designed to enhance AI inference across various chips, aiming to reduce costs and improve efficiency.
New Delhi, India Jul 8, 2026 ALN: The days of Nvidiaâs unparalleled market dominance arenât over, but challengers and choices are arising from all directions.
ZML, a hot French AI startup endorsed by Turing Award winner Yann LeCun, has released inference-performance software that allows a variety of open-source large language models to run on a variety of chips â including Nvidiaâs, AMDâs, Googleâs TPU, Apple Metal, and Intel Arc.
With ZML/LLMD, the newly launched LLM inference server, the companyâs ambition is to break existing silos and make different chips available for AI use cases at their maximum available speed, and sometimes faster, ZML founder Steeve Morin stated.
As AI becomes integrated into our work and everyday lives, optimizing inference â aka, the processing of prompts â has been outpacing model training in importance, but often feels patchy behind the scenes, with software and architecture barriers that lead to vendor lock-in, Morin explained.
The promise of achieving peak performance across a variety of chips is a technological feat, but it could also be a market disruptor, amid mounting fears over AI-related costs.
ZML hopes to provide enterprises and clouds with the option to use a mix of chips, some of which might be less costly or consume less energy. âThe idea is to give people back the power to create their own system and achieve real efficiency gains that allow [AI] to be disseminated,â Morin said.
Such a software assist may help novel AI chipmakers, many of which happen to be from Europe, Morin observed, citing Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud, and VSORA. But more than their region of origin, what matters to him is that ZML can work with them on âthings that havenât been done before anywhere in the world.â
That doesnât mean Morin is bearish on Nvidia. Heâs not, in part because of its existing supply. He mentioned that ZML has a good relationship with the AI chip giant, which has been gearing up for the rise of inference.
Inference has been an area of such intense investment that the trend has been hailed the âinference gold rush.â So ZML faces competition from companies like Baseten, recently valued at $13 billion; Inferact, from the creators of the open-source project vLLM; and RadixArk, the commercial company behind SGLang.
Both vLLM and SGLang partially compete with LLMD, but Morinâs ambitions for ZML cover a broader spectrum. âWe have reached the point where we are co-designing silicon,â he said. He credited ZMLâs lean team of 20 people as the reason why the Paris-based startup has been able to move fast, with more releases in the plans.
It also helped that this small team is well funded for its size. Thanks to his track record as VP of engineering of Zenly, which Snapchat acquired for nine figures in 2017, Morin raised $20 million from venture firms including Harry Stebbingsâ 20VC, >commit, AALVC, Drysdale Ventures, Xavier Nielâs Kima Ventures, Kindred Capital, LocalGlobe, and Puzzle Ventures.
Unlike ZMLâs first public project, the inference-focused ML framework released in 2024 and updated in March, ZML/LLMD is not open source. However, it is launching as a free product with the goal of learning about usage. âIâd rather measure and [then generate revenue] where it is most effective without hindering my growth stupidly because I have been too greedy from the get-go,â Morin said.
It is too early to tell when ZML/LLMD might become a paid product and what its adoption will look like. However, the startupâs cap table confirms that other founders are paying attention, including Dagger and Docker founder Solomon Hykes, ClĂ©ment Delangue and Julien Chaumond from Hugging Face, as well as LeCun, now with AMI Labs. This also builds the case that Europeâs AI startups can now build from home. âI couldnât do ZML anywhere but in Paris,â Morin concluded.
The emergence of startups like ZML is significant in the context of the broader AI landscape, which has traditionally been dominated by a few key players, particularly Nvidia. Nvidia has established itself as a leader in providing GPUs that are optimized for AI workloads, making it a go-to choice for many developers and enterprises looking to implement AI solutions. However, as the demand for AI capabilities grows, so does the need for diverse hardware options that can cater to various applications and budgets.
ZMLâs approach to creating software that can optimize inference across multiple hardware platforms may serve to democratize access to AI technology. By enabling developers to leverage a mix of chips, ZML could help organizations avoid vendor lock-in, which has been a significant concern for many enterprises investing in AI. This flexibility could also lead to cost savings, as companies can choose hardware that aligns with their specific needs and constraints.
The implications of ZMLâs innovation extend beyond just technical performance. As AI technology becomes increasingly embedded in various sectors, including healthcare, finance, and education, the ability to optimize inference across different chips could lead to more efficient AI applications. This could ultimately enhance productivity and drive innovation, as organizations are better equipped to deploy AI solutions that are tailored to their unique requirements.
Moreover, the competition ZML faces from other companies in the inference space underscores the growing interest and investment in AI technologies. The term "inference gold rush" highlights the urgency with which companies are trying to capitalize on the potential of AI, particularly in terms of real-time processing and responsiveness. As more players enter the market, it is likely that we will see rapid advancements in inference technologies, ultimately benefiting end-users through improved AI performance.
ZMLâs strategy of launching ZML/LLMD as a free product is also noteworthy. By prioritizing user feedback and understanding usage patterns before implementing a monetization strategy, ZML is taking a customer-centric approach that could foster loyalty and encourage widespread adoption. This strategy may also allow ZML to refine its product based on real-world applications and challenges, ensuring that it meets the needs of its users effectively.
As ZML continues to develop its offerings and build relationships with other chipmakers, its role in the AI ecosystem could evolve significantly. The collaboration with European chipmakers could lead to innovations that leverage the strengths of various hardware platforms, potentially creating a more robust and diverse AI infrastructure. This could also enhance Europeâs position in the global AI landscape, showcasing the regionâs potential to produce competitive AI technologies that can stand alongside established players.
In conclusion, ZMLâs launch of its inference-performance software marks a pivotal moment in the AI industry, with the potential to disrupt existing paradigms and challenge the dominance of major players like Nvidia. By focusing on optimizing inference across multiple hardware platforms, ZML is positioning itself as a key player in the rapidly evolving AI landscape, with the potential to drive innovation and enhance accessibility for organizations looking to leverage AI technologies. As the company navigates its growth trajectory, its impact on the AI ecosystem will be closely watched by industry stakeholders and competitors alike.
To learn more about the latest developments in Startup Sectors & Industries, stay updated with our exclusive reports and analyses on AiLensNews.