Scale without limits
Deploy compact models that eliminate GPU shortages, reduce latency, and accelerate growth.
Compressed models, the API to run them, the platform to host them and the layer that governs them.
Compressed AI and quantum-inspired optimization where compute, energy and regulation bite hardest.
Tensor networks and quantum algorithms, in production since 2019. The same mathematics that compresses our models.
Quasar 438B sets a new benchmark for European AI, outperforming comparable models on seven of eight evaluations.
Founded in 2019 by four scientists. Offices across Europe, the US and Canada. Series C targeting up to $570M.
We empower organizations to run secure, production-ready AI with tailored solutions, reducing compute costs and retaining full control across cloud, data centers and edge environments.
Who runs compressed AI
The same models, four reasons to shrink them: no GPUs to buy, a fleet to fill, racks to keep full or a device with almost no memory.
Deploy compact models that eliminate GPU shortages, reduce latency, and accelerate growth.
Run advanced AI on existing hardware with compressed models that cut CAPEX and energy use.
Increase throughput and profitability without adding racks — thanks to smaller, faster models.
Enable on-device inference with compressed models that fit limited compute and memory.
Released 2 September 2026
A reasoning model for enterprise agents and coding. It beats models with more parameters because CompactifAI removes what the network never used.
13th of 178 models worldwide, and the highest score of any European model — beating a model built with 112 billion more parameters.
curl api.compactif.ai/v1/chat/completions -d '{"model":"quasar-438b","reasoning":"high", …}'How CompactifAI works
Tensor networks are the framework physicists use for quantum many-body systems. Pointed at a weight matrix, they show which correlations carry signal and which are noise. What survives is decomposed into a chain of small tensors: the same model, a fraction of the weight.
A dense 4096×4096 weight matrix is profiled layer by layer; 66% of its correlations are found to be spurious and discarded; the remaining third is decomposed into a chain of five small tensors with bond dimension 64, then healed by a short retraining pass. The result has 95% fewer parameters and retains 97.4% of the accuracy, running 4 to 12 times faster at 50 to 80% lower cost.
The same compressed weights move from our API to your data center to a device without changing a line of your integration. Choose by where the data lives, not by where the GPUs are.
The compressed model lives at an endpoint served from Europe. Structured outputs, tool calling and always-on reasoning with effort controls.
The same weights, exported for your servers. Runs on Intel Xeon 6 CPUs and Qualcomm accelerators as well as GPUs, so capacity grows without new hardware.
Sovereign deployments for banks, utilities, defense and the public sector. Air-gapped if you need it, governed by SentinelAI either way.
Small enough for an industrial controller, a vehicle or a phone. Inference happens where the data is produced, with no round trip to a cloud.
One model. Anywhere you need it.One compressed checkpoint, exported for GPU, CPU and NPU runtimes. The same OpenAI-compatible endpoints in your rack as in our cloud. SentinelAI in front of all of it.
The AI Suite
Four products built on one idea: AI that fits the hardware, the budget and the jurisdiction you already have. Three are here; Foundry has its own section below.
Our advanced compression technology reduces LLM size, enabling faster, scalable and cost effective AI on any enterprise system or edge device.
One governed pipeline between every application and every model. Aligned with the EU AI Act, GDPR and SOC 2, adaptable to ISO, HIPAA and DORA.
Our quantum-inspired optimization platform powers the breakthroughs behind our Compact AI technology.
Foundry · early access
A unified platform for sovereign AI infrastructure: explore models, deploy GPU clusters, monitor costs and workloads, and access serverless AI services — all within a clean, intuitive control panel. Try the four steps.
https://eu-west.internal/v1/chat/completionsRepresentative panel. Cost and energy figures shown fall inside the 50–80% reduction measured by HP and Telefónica.
The system
Expertise
Training and inference costs are escalating. With model sizes exploding, efficient AI is no longer optional. Our compression technology makes large-scale AI sustainable.
Green transition
Energy-intensive infrastructures require smarter, more efficient algorithms. Compact AI reduces compute and power usage across critical applications — Telefónica measured up to 75% less energy.
Real results
The compressed models can be deployed directly on Telefónica’s network, including local facilities, making it possible to reduce energy consumption by up to 75% compared to uncompressed models.”Telefónica · network deployment
We conducted extensive technical benchmarking and were very impressed: decrease in time to first token, increase in token throughput, and models that are cheaper to run, with 50–80% cost and energy savings.”HP · benchmark programme
Model footprint cut by over 50% in a consumer support assistant, with lower latency and cost at the same response quality.Luzia · consumer assistant
Substantial reductions in power, operational cost and carbon emissions at competitive accuracy.Sopra Steria · independent evaluation
Customers and partners
Company
Founded in 2019, we recognized quantum technology’s potential to revolutionize AI and optimization solutions. This vision gave rise to Singularity, our advanced platform, created to empower clients. Guided by a customer-centric philosophy, we collaborate to tailor models for specific challenges, aiming to seamlessly integrate user-friendly, disruptive software into the industry landscape.
Tensor networks were invented to describe systems with more states than there are atoms in the universe — by keeping only the correlations that matter.
The same decomposition, applied to portfolio, grid, hydrogen and fleet problems. In production in finance since 2019, and an IBM Quantum partner.
And to the weight matrices of the largest language models. 160+ patents stand behind it.




Latest
Outperforms comparable European models on seven of eight selected Artificial Analysis evaluations.
Cryptographic agent identity and governance combined with model compression and edge deployment.
To expand the library of compressed models and scale commercial deployment worldwide.
CPUs and data-center accelerators as well as GPUs, so capacity grows without new hardware.
Compressed models deployed directly on the network, including local facilities.
After extensive technical benchmarking: faster time to first token and higher throughput.