Quasar 438B is live · AAII 43 Telefónica: up to 75% less energy Series C: targeting up to $570M

Pioneering the era of efficient & secure AI.

We empower organizations to run secure, production-ready AI with tailored solutions, reducing compute costs and retaining full control across cloud, data centers and edge environments.

Model size
100%
Accuracy retained
100%
Scroll to compress
up to 95%smaller models
4–12×faster inference
50–80%lower inference cost
Trusted by 100+ organizations

Who runs compressed AI

Built for whoever pays the compute bill.

The same models, four reasons to shrink them: no GPUs to buy, a fleet to fill, racks to keep full or a device with almost no memory.

For digital natives

Scale without limits

Deploy compact models that eliminate GPU shortages, reduce latency, and accelerate growth.

Learn more
For corporations

Extend your infrastructure

Run advanced AI on existing hardware with compressed models that cut CAPEX and energy use.

Learn more
For data centers

Maximize capacity

Increase throughput and profitability without adding racks — thanks to smaller, faster models.

Learn more
For device manufacturers

Powerful AI on any device

Enable on-device inference with compressed models that fit limited compute and memory.

Learn more

Released 2 September 2026

Europe’s highest-scoring model. Built by compressing, not by adding.

A reasoning model for enterprise agents and coding. It beats models with more parameters because CompactifAI removes what the network never used.

Quasar 438B
Parameters
438B
Output speed
176 tok/s
Context window
1M tokens
Input / output per M
$0.60 / $1.80
Artificial Analysis Intelligence Indexv4.1.1 · higher is better
  1. 43
    Quasar 438B Multiverse Computing · 438B parameters
  2. 38
    Nemotron 3 Ultra NVIDIA · 550B parameters

13th of 178 models worldwide, and the highest score of any European model — beating a model built with 112 billion more parameters.

Quasar 438B · 13th
#1 worldwide#178
Terminal-Bench 2.1 · agentic coding
0.0
AA-LCR · long-context reasoning
0.0
APIcurl api.compactif.ai/v1/chat/completions -d '{"model":"quasar-438b","reasoning":"high", …}'

How CompactifAI works

A 4096×4096 matrix, and the third of it that carries information.

Tensor networks are the framework physicists use for quantum many-body systems. Pointed at a weight matrix, they show which correlations carry signal and which are noise. What survives is decomposed into a chain of small tensors: the same model, a fraction of the weight.

A dense 4096×4096 weight matrix is profiled layer by layer; 66% of its correlations are found to be spurious and discarded; the remaining third is decomposed into a chain of five small tensors with bond dimension 64, then healed by a short retraining pass. The result has 95% fewer parameters and retains 97.4% of the accuracy, running 4 to 12 times faster at 50 to 80% lower cost.

One model. Four places to run it.

The same compressed weights move from our API to your data center to a device without changing a line of your integration. Choose by where the data lives, not by where the GPUs are.

01Cloud API
01 · Cloud API

Call it. One key, OpenAI-compatible.

The compressed model lives at an endpoint served from Europe. Structured outputs, tool calling and always-on reasoning with effort controls.

TTFT 0.31 s · 176 tok/s
02 · Data center

Rack it. More tenants per node.

The same weights, exported for your servers. Runs on Intel Xeon 6 CPUs and Qualcomm accelerators as well as GPUs, so capacity grows without new hardware.

4–12× faster inference per node
03 · Edge and on-prem

Keep it inside. Data never leaves.

Sovereign deployments for banks, utilities, defense and the public sector. Air-gapped if you need it, governed by SentinelAI either way.

up to 75% less energy, measured by Telefónica
04 · On device

Ship it. It fits in the memory you have.

Small enough for an industrial controller, a vehicle or a phone. Inference happens where the data is produced, with no round trip to a cloud.

0 bytes to the cloud

One model. Anywhere you need it.One compressed checkpoint, exported for GPU, CPU and NPU runtimes. The same OpenAI-compatible endpoints in your rack as in our cloud. SentinelAI in front of all of it.

The AI Suite

Compress it. Run it. Host it. Govern it.

Four products built on one idea: AI that fits the hardware, the budget and the jurisdiction you already have. Three are here; Foundry has its own section below.

Deploy compressed AI models anywhere.

Our advanced compression technology reduces LLM size, enabling faster, scalable and cost effective AI on any enterprise system or edge device.

  • 50–80% lower inference costs
  • Up to 2× faster inference
  • Close to 100% accuracy retention
Explore CompactifAI
QUASARMISTRALLLAMADEEPSEEK

Every request inspected, every response audited.

One governed pipeline between every application and every model. Aligned with the EU AI Act, GDPR and SOC 2, adaptable to ISO, HIPAA and DORA.

  • Redacts PII and secrets before they leave
  • Screens prompt injection and policy breaches
  • Observes spend and risk per call
  • Audits every request and response
See SentinelAI
QUASARLLAMA

Optimizing complex problems with Singularity.

Our quantum-inspired optimization platform powers the breakthroughs behind our Compact AI technology.

  • Portfolio optimization, also inside Excel
  • Grid, hydrogen and fleet optimization
  • IBM Quantum partner
Explore Singularity

Foundry · early access

Pick a model. Pick where it runs. Watch what it costs.

A unified platform for sovereign AI infrastructure: explore models, deploy GPU clusters, monitor costs and workloads, and access serverless AI services — all within a clean, intuitive control panel. Try the four steps.

Foundry eu-west · Donostia 01 / 04 · catalogue early access
Compressed model catalogue4 of 40
  • Quasar 438Breasoning · agents · coding438B43 AAII
  • Llama 3.3 70Bcompressed by CompactifAI70Bready
  • Mistral Largecompressed by CompactifAI123Bready
  • DeepSeek R1compressed by CompactifAI671Bready

Representative panel. Cost and energy figures shown fall inside the 50–80% reduction measured by HP and Telefónica.

Expertise

Trusted by leading efficiency-driven industries.

Training and inference costs are escalating. With model sizes exploding, efficient AI is no longer optional. Our compression technology makes large-scale AI sustainable.

Green transition

Less compute is less power.

Energy-intensive infrastructures require smarter, more efficient algorithms. Compact AI reduces compute and power usage across critical applications — Telefónica measured up to 75% less energy.

Your sector is not on this list? Wherever compute, energy or regulation is the constraint, the answer is the same shape. Tell us which one is yours. Talk to an engineer
FinanceSovereign LLMs, portfolio and risk optimization
EnergyGrid, storage and demand optimization
ManufacturingEdge inference on the line
Health & Life SciencesPrivate models on hospital hardware
EngineeringDigital twins and simulation
AerospaceFlight-operations optimization
CybersecurityGoverned, air-gapped inference
DefenseSecure-by-design AI chips
ChemistryMolecular and process modeling
HydrogenStorage and transport optimization

Real results

Validated by global leaders.

up to 75% less energy per deployment, running compressed models directly on the network.
The compressed models can be deployed directly on Telefónica’s network, including local facilities, making it possible to reduce energy consumption by up to 75% compared to uncompressed models.”Telefónica · network deployment
50–80% lower cost and energy measured across an extensive internal benchmarking programme.
We conducted extensive technical benchmarking and were very impressed: decrease in time to first token, increase in token throughput, and models that are cheaper to run, with 50–80% cost and energy savings.”HP · benchmark programme
over 50% smaller footprint in a consumer support assistant, at the same response quality.
Model footprint cut by over 50% in a consumer support assistant, with lower latency and cost at the same response quality.Luzia · consumer assistant
independently evaluated by a third party, not by us.
Substantial reductions in power, operational cost and carbon emissions at competitive accuracy.Sopra Steria · independent evaluation

Customers and partners

Telefónica
Bosch
Bank of Canada
HP
Intel
AWS
IBM Quantum
BASF
Renault
Leonardo
Baker Hughes
Crédit Agricole
Deloitte
PwC
EY
Indra
Navantia
OVHcloud
Cerebrium
Luzia
Sopra Steria
BDR
European Union
Real Sociedad

Company

The biggest quantum software provider in Europe.

Founded in 2019, we recognized quantum technology’s potential to revolutionize AI and optimization solutions. This vision gave rise to Singularity, our advanced platform, created to empower clients. Guided by a customer-centric philosophy, we collaborate to tailor models for specific challenges, aiming to seamlessly integrate user-friendly, disruptive software into the industry landscape.

  1. Quantum many-body physics

    Tensor networks were invented to describe systems with more states than there are atoms in the universe — by keeping only the correlations that matter.

  2. Singularity · optimization

    The same decomposition, applied to portfolio, grid, hydrogen and fleet problems. In production in finance since 2019, and an IBM Quantum partner.

  3. CompactifAI · model compression

    And to the weight matrices of the largest language models. 160+ patents stand behind it.

Patents
160+
Customers
100+
Series C target, July 2026
$570M
Offices
Europe · US · Canada
Enrique Lizaso Olmos, painted portrait
Enrique Lizaso OlmosCo-founder and CEO
Román Orús, painted portrait
Román OrúsCo-founder and Chief Scientific Officer
Samuel Mugel, painted portrait
Samuel MugelCo-founder and CTO
Alfonso Rubio-Manzanares, painted portrait
Alfonso Rubio-ManzanaresCo-founder

Latest

  • Launch ·

    Introducing Quasar 438B, Europe’s leading AI model

    Outperforms comparable European models on seven of eight selected Artificial Analysis evaluations.

  • Partnership ·

    FIOR Group partnership for secure edge and sovereign AI

    Cryptographic agent identity and governance combined with model compression and edge deployment.

  • Funding ·

    Series C announced, targeting up to $570M

    To expand the library of compressed models and scale commercial deployment worldwide.