If the name PrismML has not yet cemented itself in your daily lexicon of artificial intelligence developments, it is time to take notice. While the startup has not dominated headlines with massive venture capital hauls—having secured a modest $22.25 million seed round to date—it is rapidly becoming a focal point for industry insiders. The reason is simple: PrismML is architecting a potential paradigm shift in how large language models (LLMs) are deployed, moving away from the "bigger is better" philosophy that has defined the current AI arms race. At its core, PrismML is betting that high-performing, sophisticated reasoning models do not actually need to be "large" in the traditional, resource-intensive sense. The startup is effectively shrinking these reasoning powerhouses down to a footprint that allows them to run locally on standard personal computers and even high-end smartphones. This capability has not gone unnoticed by the industry’s heavyweights, with persistent rumors suggesting that the company is currently in talks with Apple—a move that would align perfectly with the tech giant’s strategy of integrating more robust, on-device AI into its hardware ecosystem. While CEO Babak Hassibi declined to comment on those specific rumors, the implications of such a partnership are clear: the future of AI may not live in a massive server farm, but in the palm of your hand. The latest milestone in this effort arrived this Thursday with the release of Bonsai 2 27B. This model represents the cutting edge of the company’s compression research. By taking Qwen3.8 27B—a highly regarded open-source model developed by Alibaba—and applying its proprietary compression techniques, PrismML has managed to shrink the model down to a mere 5.9 GB. To put this in perspective, this represents a 9x to 10x reduction in memory requirements compared to the original, making it small enough to run on consumer-grade hardware without sacrificing the heavy lifting required for complex reasoning tasks. The intellectual backbone of PrismML is deeply rooted in academic rigor. Founded by a cohort of researchers from the California Institute of Technology (Caltech), the startup is led by Babak Hassibi, a Caltech professor and a recognized expert in signal processing and compression technologies. The company’s trajectory is further bolstered by the involvement of Ion Stoica, a luminary in the computing space who serves as an adviser. Stoica is perhaps best known as a co-founder of Databricks and the director of the University of California, Berkeley’s esteemed Sky Computing Lab. The lab has established itself as a prolific engine for innovation, having played a formative role in the development of various influential startups and technologies, including Letta and the inference-focused SGLang. With backing from marquee investors like Khosla Ventures, Cerberus Capital, and Caltech itself, PrismML is well-positioned to navigate the challenging landscape of AI infrastructure. PrismML is certainly not the only entity exploring the frontiers of model compression. The race to make AI more efficient is heating up, with companies like Multiverse Computing, founded by a distinguished professor from Spain’s Donostia International Physics Center, also making significant strides in the field. Unlike PrismML, which has kept its funding relatively lean, Multiverse Computing has pursued a strategy of raising substantial capital to push its compressed models into the mainstream. However, Hassibi contends that PrismML’s approach offers a unique advantage: the preservation of performance. According to the company, the Bonsai 2 model achieves this high level of compression while losing virtually no performance compared to its original, uncompressed counterpart. The model currently matches 98% of the aggregate benchmark scores of the original Qwen model. This is a marked improvement from the first generation of Bonsai, released in March, which matched 95% of its predecessor’s performance. The traction for this approach is evident in the adoption numbers; the original Bonsai model has already surpassed 11 million downloads, and the company’s smaller, specialized variants have seen an additional 2.6 million downloads. This iterative progress suggests that PrismML is closing the gap between compressed and full-sized models. Whether the company can reach 100% parity remains an open question—one that even Hassibi acknowledges is difficult to answer definitively. "Compression will likely always have some impact," he notes. However, from a practical standpoint, the search for "perfect" benchmark parity may be more academic than necessary. Given that even uncompressed LLMs are not inherently flawless and that current benchmarks often fail to perfectly capture the nuance of real-world performance, a 2% degradation is unlikely to meaningfully alter the user experience in day-to-day tasks. Moreover, as recent developments in the industry have shown, the software harness surrounding an AI model is often just as critical as the model itself when it comes to delivering accurate, reliable outputs. The secret to PrismML’s success lies in its innovative approach to weight compression. In the context of neural networks, "weights" are the fundamental building blocks of intelligence—the numerical values that a model learns and stores during the training process to inform its decision-making. In a standard model, each weight typically requires 16 bits of data. PrismML utilizes a technique known as "ternary" weights, which simplifies the information stored for each weight down to just three possible values: +1, -1, or 0. By drastically reducing the precision required to store these weights, the total memory footprint of the model is slashed, allowing for massive efficiency gains without stripping away the model’s reasoning capabilities. For those interested in the technical minutiae, the project’s documentation on Hugging Face offers a comprehensive look at the mathematical underpinnings of this compression methodology. Looking ahead, PrismML has set its sights on scaling its technology. The company’s immediate goal is to apply this ternary weight compression to even larger models. "The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there," Hassibi said. He explains that as the size of the base model grows, there is effectively more "room" to compress the data without compromising the core intelligence of the system. In his view, the general trend points toward larger models becoming the ideal candidates for this level of extreme compression, making the pursuit of 100% performance parity a more achievable target as the models grow in complexity. For advisers like Ion Stoica, the promise of this technology is not just about the technical achievement, but about the democratization of high-end AI. "You are going to have intelligence at your fingertips, and it’s going to be free because it’s going to run on the device you already bought," Stoica explains. Beyond the economic benefits of removing the need for costly cloud-based inference, there is the crucial element of privacy. By shifting the processing burden from the cloud to the local hardware, users no longer need to send their data to a remote server, ensuring that sensitive information remains contained within the device itself. As the AI industry continues to grapple with the unsustainable costs and energy demands of massive, cloud-bound models, PrismML’s work represents a critical path toward a more sustainable and accessible future. By proving that intelligence can be compressed without being diminished, the startup is not just refining an existing technology; it is rewriting the rules of deployment, bringing the next generation of reasoning models to the devices already sitting on our desks and in our pockets. With the promise of larger, more capable models on the horizon, PrismML’s trajectory suggests that the most significant breakthroughs in AI may not come from building bigger models, but from making the best ones fit anywhere. Post navigation The Ghost in the Machine: Are AI Models Developing an Instinct for Self-Preservation?