Watch the swarm work.
Then measure it.
AgentSaaSy is an R&D platform for enterprise AI Agent stacks. A test harness orchestrates AI Agents through real workloads while AEQ, the Agent Efficiency Quotient, scores the architecture itself: how much business value the design delivers per token it consumes. Specs do not predict adequacy. Only measuring the model-workload pair does.
The Swarm, Live
A simple workflow, end to end: a request enters, the harness routes it to a planner, fans work out to specialist AI Agents in parallel, a cross-family judge reviews the result, and the AEQ gate issues a pre-registered verdict before anything ships. Every hop is metered.
AEQ Monitor (simulated)
--BVD units / 1K tokensRun Counters
Harness Log
Architecture Layers
Four layers do the work. One measurement plane cuts across all of them. Hover or tap a layer to inspect it.
AgentSaaSy Application Layer
Workflow definitions, business value rubrics, and the SaaS-substitution surface where AI Agent stacks replace seat-licensed software.
Harness & Orchestration Layer
The conductor. Routes requests, fans out parallel work, enforces pre-registered gate thresholds, and coordinates cross-family judging.
Agent Layer
The swarm: planner, retriever, workers, and an independent judge from a different model family than the workers it reviews.
Model Layer
Frontier APIs and quantized local models, treated as interchangeable capacity. The pair (model + workload) is what gets measured, never the spec sheet.
Layer Detail
The Stack
What actually runs underneath the demo above.
From Simple to Swarm-Scale
This page shows one workflow. The architecture is built to grow.
One workflow, fully metered
Single request pipeline with parallel fan-out, cross-family judging, and a pre-registered AEQ gate on every run.
Multi-workflow swarms
Concurrent workflows sharing the agent pool, tiered model routing, and per-workflow AEQ baselines to catch architecture waste.
AEQ Grid certification
The full 3x3x3 grid: query classes by model tiers by repeated runs, producing GREEN / YELLOW / RED verdicts for model-workload pairs.
The measured version
Everything above this line runs on simulated data. The same architecture was put through a real workload and the result was written down: an enterprise asset management build answering at $0.0009 per query, with the run date and the pricing caveat on the record.
The enterprise asset management case studyThe technical white paper behind itThe AEQ specification the score is defined by