Alchemy
Alchemy is a platform that hosts a squad of specialized AI assistants for software engineering — used by analysts, developers, testers, product owners, and project managers to move faster with fewer mistakes. The name fits: alchemy is the art of fusing separate elements into something more valuable than either alone, and this platform is built as an alloy of natural and artificial talent — human expertise and AI capability, combined rather than substituted. It automates the tedious parts of the job and strengthens decision-making with AI, while keeping humans in control of every outcome.

Human oversight of AI-generated work
Generic AI tools promise speed but often trade it for risk: over-reliance without oversight, loss of team context and domain expertise, disconnection from business reality, and unvalidated output that quietly makes its way into production.
Alchemy is built on the opposite premise: AI amplifies human judgment, it doesn't replace it. Humans keep creative control and make the final call; assistants handle repetitive work, draft first passes, and get better from the feedback they receive.
Capabilities of the platform
Persisted workspaces for chats, documents, and collaborative editing
Every conversation and artifact is persisted and organized, so work survives the session it happened in.
Assistants read and extract information from your own material instead of a blank page.
Several people can edit the same artifact at once, alongside the assistant, with comments and feedback on every suggestion.
A flexible, observable web platform for organization-wide use
A modern web interface is included out of the box, with the account and usage controls a whole team needs.
The platform is parametric with respect to both LLM and provider — no lock-in to a single model or vendor.
Advanced LLM tracing gives real observability into what an assistant did and why, instead of a black box.
The chat and the artifact appear side by side

Human involvement enforced by design
Every assistant built on Alchemy follows the same principle, enforced by design rather than left to good intentions:
Checkpoints, audit trails, and feedback built into every interaction
Assistants draft, analyze, and suggest; people review, correct, and decide.
Phrasing like “what if” or “suppose” keeps an artifact untouched and returns a proposal instead of an edit — brainstorming without side effects.
Every suggestion can be rated and annotated, supporting collaborative work and continuous improvement.
Every AI change is reviewable before it lands

That’s what keeps the output usable — no unreviewed AI slop landing in your specs, your tests, or your codebase, and no loss of your team’s context, domain expertise, or ownership of what gets shipped.
Acquiring and customizing assistants
Assistants available from our catalog or built to order
Alchemy already powers assistants we’ve built ourselves — Cuprum for use case modeling, Iridium for knowledge management, and the assistants inside CuTE for test automation.
When nothing in the catalog fits, we build an assistant specifically for your process and systems — on the same platform, with the same human-in-the-loop guarantees.
Every assistant, ours or bespoke, shares the same foundation, so adding one more never means starting from scratch.
Customization of the task, its oversight, and its output
The task itself — and the best practice it applies — is adapted to your established processes and approval workflows.
Review processes, confidence thresholds, and audit trails match how your team already governs its own work.
Output takes your document templates, naming conventions, and terminology — so it looks like it came from your team, because it did.
Interface and infrastructure options
An assistant is reached through the built-in web UI, or wired into another one entirely — a customer’s own portal, an IDE such as VSCode, or a modeling tool such as Enterprise Architect or Miro.
Deployment is just as flexible: Alchemy can run on site, inside your own VPC, for organizations that need data and LLM traffic to stay within their own infrastructure and security perimeter.
Further automation across the software lifecycle
The same pattern extends to other steps of the lifecycle — automatically breaking use cases into well-sliced user stories, generating acceptance criteria and Gherkin scenarios, mapping feature impact, standardizing bug reports, documenting code, or scoring test suite quality.
Each is a candidate for an off-the-shelf assistant or a bespoke one, built the same way.
Observability and evaluation through the Control Hub
Why it matters: across every assistant running on Alchemy, the Control Hub turns “we think it got better” into objective evidence you can act on directly — catching regressions before they reach production, tracing a bad answer to the exact step that caused it, and keeping the whole cluster of assistants improving against how they’re actually used. Measurable quality — for one assistant or for all of them at once.
Real-time, logged observability of agent activity
Every assistant on Alchemy can be connected to Alchemy’s Control Hub, which provides live observability into what each agent is doing, step by step, turning “the output was wrong” into “here’s exactly why.”
Execution parameters and key events are captured for every run in a full, statistically verifiable log.
Structured evaluations derived from production usage data
AI engineers run standardized, auditable evaluations against selected assistants or LLMs, on demand — replacing ad hoc spot checks with evidence comparable across runs, models, and versions.
Results are reviewed, tagged, and preserved in a searchable knowledge base, shareable with reviewers.
Feedback from production use feeds new evaluation datasets, and can be turned automatically into evaluations checked against AI-as-a-judge and LLM observability techniques.
Project funding
Alchemy is co-financed by the European Union and the Friuli Venezia Giulia Region under the PR FESR FVG 21-27 program. Admitted expenditure is €115,630.20, with a €80,941.14 contribution (40% EU-funded).
Read the project summary card (PDF)
