{"id":263610,"date":"2026-09-17T22:34:09","date_gmt":"2026-09-17T22:34:09","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/17\/prismml-hopes-its-tiny-llm-will-change-how-we-all-use-ai-techcrunch\/"},"modified":"2026-09-17T22:34:09","modified_gmt":"2026-09-17T22:34:09","slug":"prismml-hopes-its-tiny-llm-will-change-how-we-all-use-ai-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/17\/prismml-hopes-its-tiny-llm-will-change-how-we-all-use-ai-techcrunch\/","title":{"rendered":"PrismML hopes its tiny LLM will change how we all use AI | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">If AI lab <a rel=\"nofollow noopener\" href=\"https:\/\/prismml.com\/\" target=\"_blank\">PrismML<\/a> isn\u2019t on your radar yet, it should be \u2014 not because it\u2019s raised gobs of money (it hasn\u2019t yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it\u2019s developing.<\/p>\n<p class=\"wp-block-paragraph\">PrismML is betting that capable, high-performing, reasoning large language models don\u2019t, in fact, have to be large.<\/p>\n<p class=\"wp-block-paragraph\">It is making reasoning models so small they can fit on PCs and smartphones. (It\u2019s even rumored <a rel=\"nofollow noopener\" href=\"https:\/\/www.cnbc.com\/2026\/07\/14\/apple-prismml-ai-compression-iphone.html\" target=\"_blank\">to be in talks with Apple<\/a>, though CEO Babak Hassibi declined to comment on that to TechCrunch.)<\/p>\n<p class=\"wp-block-paragraph\">On Thursday, PrismML <a rel=\"nofollow noopener\" href=\"https:\/\/prismml.com\/news\/bonsai-2-27b\" target=\"_blank\">released Bonsai 2 27B<\/a>, its latest in a family of models, which compresses Qwen3.8 27B, a widely used open-source model from Alibaba, down to 5.9 GB. That\u2019s small enough to fit on a PC and, possibly, a high-end smartphone. It\u2019s a 9x to 10x reduction in memory versus the original.<\/p>\n<p class=\"wp-block-paragraph\">PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies. The startup also counts Ion Stoica as an advisor. Stoica is a co-founder of Databricks (and other companies) and the director of Berkeley\u2019s famed Sky Computing Lab, which has birthed many technologies and startups, from <a href=\"https:\/\/techcrunch.com\/2024\/09\/23\/letta-one-of-uc-berkeleys-most-anticipated-ai-startups-has-just-come-out-of-stealth\/\" target=\"_blank\" rel=\"noopener\">Letta<\/a> to <a href=\"https:\/\/techcrunch.com\/2026\/01\/21\/sources-project-sglang-spins-out-as-radixark-with-400m-valuation-as-inference-market-explodes\/\" target=\"_blank\" rel=\"noopener\">SGLang<\/a>. <\/p>\n<p class=\"wp-block-paragraph\">PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech.<\/p>\n<p class=\"wp-block-paragraph\">This startup is certainly not the only company working on LLM compression tech. Multiverse Computing, founded by a well-known professor from Spain\u2019s Donostia International Physics Center, is another. (And Multiverse Computing <a href=\"https:\/\/techcrunch.com\/2026\/03\/19\/multiverse-computing-pushes-its-compressed-ai-models-into-the-mainstream\/\" target=\"_blank\" rel=\"noopener\">has raised gobs of cash<\/a>.)<\/p>\n<p class=\"wp-block-paragraph\">But Hassibi says that PrismML\u2019s compression tech is unique because its LLMs have lost virtually no performance compared with the originals. Bonsai 2 matches 98% of Qwen\u2019s aggregate benchmark scores. That\u2019s up from the first Bonsai, released a couple of months ago in March, that matched 95%. That original model has already been downloaded over 11 million times, and PrismML\u2019s even smaller models have been downloaded another 2.6 million times, the company says.<\/p>\n<p class=\"wp-block-paragraph\">So this shows that PrismML\u2019s compression results have improved from one release to the next. Whether it could ever get to 100% benchmark performance parity is a question that remains to be seen. Compression will likely always have <em>some<\/em> impact, Hassibi says.<\/p>\n<p class=\"wp-block-paragraph\">Still, perfect benchmark parity is fairly academic anyway. LLMs are not so accurate in their uncompressed form, and benchmarks not so perfectly reflective of actual tasks, that a 2% degradation would likely meaningfully affect how a model performs in actual use. (Plus, the surrounding software \u2014 the harness a model runs inside of \u2014 <a href=\"https:\/\/techcrunch.com\/2026\/08\/21\/nvidia-just-showed-that-the-harness-not-the-ai-model-is-now-the-real-hero\/\" target=\"_blank\" rel=\"noopener\">matters a lot when it comes to accuracy<\/a>, too.)<\/p>\n<p class=\"wp-block-paragraph\">PrismML says it achieves this by shrinking the \u201cweights\u201d that make up a model \u2014 weights are, essentially, the information a model learns and stores during training. Normally, each weight requires 16 bits. PrismML\u2019s approach, called \u201cternary\u201d weights, simplifies that down to three: +1, \u22121, or 0. With far smaller values to store for each weight, the model takes up dramatically less space. (For a deeper dive on the compression technique, here\u2019s the project\u2019s <a rel=\"nofollow noopener\" href=\"https:\/\/github.com\/fpgasystems\/ternaryLLM\" target=\"_blank\">GitHub page<\/a>.)<\/p>\n<p class=\"wp-block-paragraph\">The startup\u2019s next goal is to apply this compression technique to even bigger models. \u201cThe next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there,\u201d Hassibi told TechCrunch.<\/p>\n<p class=\"wp-block-paragraph\">As model size grows, he added, \u201cThere is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, it\u2019s easier to get to 100%.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Stoica tells us that he\u2019s excited for this tech because it\u2019s making it possible for advanced models to run on users\u2019 devices. \u201cYou are going to have intelligence at your fingertips, and it\u2019s going to be free because it\u2019s going to run on the device you already bought. It\u2019s also going to be private, because you\u2019re not going to send it to the cloud.\u201d<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/09\/17\/prismml-hopes-its-tiny-llm-could-change-how-we-all-use-ai\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>If AI lab PrismML isn\u2019t on your radar yet, it should be \u2014 not because it\u2019s raised gobs of money (it hasn\u2019t yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it\u2019s developing. PrismML is betting that capable, high-performing, reasoning large language models don\u2019t, in [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":263611,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-263610","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263610","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=263610"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263610\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/263611"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=263610"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=263610"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=263610"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}