Skip to content
uk-ai.news

Sqwish raises £1.7M to squeeze the cost out of running AI

By Investment Source: Sqwish

Sqwish, a Cambridge startup, has raised seed funding, reported at around £1.7M, for software aimed at one of the least glamorous but most pressing problems in AI: how much it costs to run. The company was founded by Ushnish Sengupta, who did a Cambridge engineering PhD in Bayesian machine learning, and Federica Freddi, also a Cambridge engineer, and it came out of the university’s START accelerator. Its backers include Cambridge Enterprise Ventures, Vento, and Founders at the University of Cambridge.

What Sqwish builds is a layer that sits between an application and the language model it calls. In real time, it compresses the prompt and context being sent to the model, by up to tenfold according to the company, so each request uses fewer tokens while the answer stays as good. Since models are billed by the token, fewer tokens means a smaller bill and faster responses. On top of that, a reinforcement-learning engine tunes which model and how much context to use against real business outcomes, such as conversions, rather than just technical measures like latency. The pitch is aimed squarely at the “agent economy”, the growing set of production apps built on AI agents where the cost of all those model calls adds up fast.

It is worth reading alongside a much larger round we covered this week. Where Callosum attacks the cost of AI at the level of the chip, routing each workload to the hardware best suited to it, Sqwish attacks it at the level of the prompt. Different layer, same underlying bet: AI is expensive to run, and a cluster of British companies is forming to make it cheaper, from the silicon all the way up to the tokens. That Sqwish is a Cambridge spin-out adds to the university’s steady output of AI-infrastructure startups.

The reason this category is having a moment is simple arithmetic. Every call to a model is charged by the token, and for a busy application those inference costs quickly dwarf almost everything else. Trimming the tokens per call is one of the most direct levers a company has to cut its AI bill, which is why an optimisation layer like this, unglamorous as it sounds, is exactly the sort of thing enterprises are starting to look for.

Read the original story on Sqwish .

Companies in this story