Bring your own model key

Customer AI that earns its next version.

UnifyAI turns everyday sales, marketing and support work into a measured loop. Run a task, correct the answer once, and let the correction compete for the right to become your default — on examples it has never seen.

API keys never persisted Held-out evaluation One-click rollback

The loop

Four steps, and none of them are a thumbs-up.

Most AI tools collect ratings and hope. UnifyAI collects the corrected answer, turns it into a candidate prompt, and refuses to ship it until it beats what you already have.

01

Run a real task

Bring the account, campaign or ticket context you already have. The active prompt and model for that team produce a draft.

Sales, marketing and support each keep their own prompt.
02

Correct it once

Approve the draft, or rewrite it the way it should have been written. The correction is the training signal — not a thumbs-up.

Six reviewed tasks unlock an experiment.
03

Evaluate a candidate

Your corrections propose a revised prompt. It is scored against the current one on held-out examples it has never seen.

A frozen rubric scores both sides identically.
04

Promote or discard

Read the side-by-side evidence and promote only if the candidate clears the gate. Every promotion is reversible.

Rollback restores the previous prompt and model.

The promotion gate

A new prompt ships only when the evidence says so.

The candidate and the current prompt answer the same held-out tasks and are graded by the same frozen rubric. All four conditions must hold — there is no override button.

You see both answers before deciding

Baseline and candidate outputs sit side by side with their scores and token counts.

Nothing is one-way

Rollback restores the previous prompt and model for that team, without touching completed tasks.

Promotion conditionsAll must pass
Scores at least 80Grounding, task completion, clarity and next step, on a rubric fixed before the run.
score >= 80
Never regressesA candidate that scores below today's prompt cannot be promoted, however good it reads.
candidate >= base
Wins or costs lessA tie only passes when the candidate uses fewer tokens for the same quality.
better || cheaper
Passes a safety checkInvented commercial promises, ungrounded policy claims or asserted actions fail outright.
safe == true

Security & data

Built so you can put real customer context in it.

Your model key never lands in our database

The key lives in the current browser tab and is sent over HTTPS for model calls only. It is not written to storage or logs.

Records are scoped server-side

Every read and write is bound to your authenticated identity. Mutations check request origin, use prepared SQL and verify record revisions.

Your feedback trains nothing shared

Learning updates the prompt and model selection inside your own workspace. No shared weights, no cross-account mixing.

Every change leaves a trace

Runs, reviews, evaluations and promotions are recorded so you can see why the current prompt is the current prompt. Read the privacy & data page.

This release

What it does, and what it does not.

The product refuses to invent claims about your customers. It would be strange to invent them about itself.

Available today

  • Individual workspaces with private task history
  • Sales, marketing and support prompts, tracked separately
  • Durable, resumable evaluation runs
  • Side-by-side evidence before any promotion
  • One-click rollback to the previous strategy
  • Email or ChatGPT sign-in, and a workspace export

Not in this release

  • Shared team memberships and role permissions
  • CRM OAuth connectors — context is entered directly
  • Model-weight fine-tuning — prompt and model only
  • Usage billing — your provider bills you for model use

Questions

Before you put real work in it.

Do I need my own model API key?

Yes. UnifyAI calls the provider you choose — UnifyAPI or OpenAI — with a key you supply per browser tab. Your provider bills you directly for model usage, and their bill is the authoritative one.

How many tasks before it actually improves?

An experiment needs at least six reviewed tasks in one team. The earlier ones become training examples; the last two are held out so the candidate is scored on work it has never seen.

What happens if a new prompt turns out worse in practice?

Roll it back. Promotion stores the previous prompt and model, and rollback restores them for that team. Tasks you already ran are untouched.

Is my customer data used to train a model?

No. Optimization happens at the prompt and model-selection layer inside your workspace. Nothing is pooled across accounts and no model weights are updated. Your provider's own data-handling terms still apply to what you send them.

Does it fine-tune the model?

No — and that is deliberate. Prompt and model changes are inspectable, cheap to evaluate and instantly reversible, which is what makes the promotion gate meaningful.

Start with one real task

The first version is never
the one worth keeping.

Create a workspace, run the task you were going to write by hand anyway, and keep the correction you would have thrown away.