24 Aug 2026

Subconscious Raises $5.1 Million to Build the Inference Platform for Long-Running Agents

Stand: Start Up Zone
Jack O'Brien

CAMBRIDGE, MA - August 21, 2026 - Subconscious is announcing $5.1 million in funds raised to launch an inference platform designed for long-running agents. MassVentures led the round, with participation from Foothill Ventures, Underscore VC, E14 Fund, and the Agent Fund among others. The inference platform is available today for developers, with an on-prem deployment package available for enterprises.

Language models are used in many ways: as a chatbot, as a classifier, as a document writer, or as an AI agent. Among all these uses, agents are extremely computationally intensive and require processing thousands of times as many tokens over long periods of time. Agents are the most expensive way to use AI models, but they generate the most valuable work. Already they’ve transformed software engineering, and they’re growing in usage among salespeople, lawyers, scientists, marketers, consultants, and all kinds of knowledge work. Subconscious believes agents will make up virtually all inference in the very near future.

Subconscious built an inference platform to take advantage of the unique challenges in powering long-running agents. Born out of MIT research into inference, the system uses dynamic context compression and highly efficient caching for a stepwise gain in performance. For agents that consume beyond 200k tokens, Subconscious can generate tokens faster, extend the effective context window of models to 5m+ tokens, improve **performance on key agentic benchmarks like coding and workflows, and decrease costs by up to 80%.

On the TriE benchmark, a benchmark built to measure system performance, Subconscious’s inference runtime generated tokens 3.5 times faster than SGLang on coding tasks and supported 2.3x as many concurrent requests. Handling more concurrent requests means squeezing more work out of GPUs, effectively doubling or tripling the size of a cluster.

On DeepSWE, a benchmark built around long coding tasks, GLM 5.2 hosted on Subconscious solved 46 percent of problems at an average cost of $2.79. The same model served on standard inference infrastructure scored 44 percent and cost $3.92. Compression and caching from Subconscious improve both the model's efficiency and its accuracy.

Subconscious is live in production today powering coding agents and agentic products. In July, a 20-person engineering team switched from using Claude to using the GLM 5.2 model hosted on Subconscious to power their coding agents. In the two months since, they’ve cut their monthly AI spend from $40k to $6k, their engineers report faster token throughput, and they have yet to hit their rate limits. One engineer on the team ran an extremely long agent trace across 4,571 turns and 9,556 tool calls, and Subconscious recorded 449 million tokens where a conventional runtime would have billed 2.6 billion. Just as important is what the engineering team didn’t report: despite aggressive context compression, they reported no loss in model capability.

As AI costs have skyrocketed many companies have moved to cut costs, but engineers want more agents running for longer periods of time on harder problems. Subconscious allows companies to have it all.

"Open models finally got good enough this summer that their quality vs closed source models stopped being a compromise," said Jack O'Brien, Co-Founder and CEO of Subconscious. "Meanwhile every engineering team I talk to has put a ceiling on what it spends per developer. Teams need to spend less, but engineers are addicted and there’s no going back. We run open models in a way that's enhanced for coding agents, so teams can have it all."

"The Subconscious team is the best team on the planet to solve one of AI's toughest challenges," said Stacy Swider, VP of Investments at MassVentures. "We could not be more excited to be a part of their growth as they scale up to support thousands of companies and billions or even trillions of AI agents."

The Subconscious inference platform is live today, and teams can start powering coding agents like Claude Code, Codex, Pi, and OpenCode with Subconscious in about 30 seconds. Teams can also deploy the inference system on their own GPUs for maximum cost savings, fully air-gapped data protection, and ultimate control.

About Subconscious

Subconscious is an inference platform designed for long-horizon agents. Engineers choose Subconscious to power their agents to finish complex tasks faster, run for longer, and lower their costs substantially. The company is based in Cambridge, Massachusetts, and has raised $5.1 million led by MassVentures.

Subconscious aims to build the necessary infrastructure to power trillions of agents. Learn more at subconscious.dev.

Media Contact
Press Team, Subconscious
press@subconscious.dev

Loading