Email that leaves a paper trail
Own the send path, the receive path, and the agent in between.
RelayFlow is a self-hosted transactional email platform. A REST send API with its own TypeScript SDK, a real inbound pipeline with SPF, DKIM and DMARC verdicts, and a runtime for putting an LLM agent behind a mailbox — on your hardware, against your Postgres.
Agent run
Completed- Mailbox
- newsletters@
- Model
- claude-haiku-4-5
- Tokens
- 914
- Cost
- $0.00133
- LLM latency
- 1,882 ms
- store_summary
- executed
- Violations
- none
A measured example from our own environment — not a guarantee, an average, or a benchmark.
Who this is for
Reasons to own this instead of renting it.
- Your customers' mail cannot sit on a vendor's servers.
- Data residency, a regulated dataset, or a contract that names where the bytes live.
- You want an agent reading a mailbox without handing that mailbox to a third party.
- The mailbox, the rules and the run records stay in your deployment; the model call you configured is the only thing that leaves it.
- You need an audit record you control.
- In your database, on your retention schedule, readable with SQL.
- You are deploying where a hosted API cannot reach.
- A private network, a customer's own account, or somewhere egress to someone else's SaaS is not something you can ask for.
- You are paying per message at a volume where owning it is cheaper.
- Throughput here is a database and a worker, not a plan tier.
The tradeoff
Every hosted email API rents you a slice of someone else's infrastructure.
Their servers
Your customers' mail sits on hardware you do not control, in a jurisdiction you did not choose, reachable by staff you have never met.
Their retention policy
How long message bodies live, and who can read them, is set by someone else's product roadmap rather than your data agreement.
Their rate limits
Throughput is a plan tier. Volume that would be cheap to run yourself becomes a per-message line item that grows with you.
No real place for an agent
You can forward a webhook to a script. What you cannot do is put a governed agent on the receive path, with an action allow-list and an audit record you own.
RelayFlow gives you the whole post office. And because you own the receive path, a mailbox stops being a dead drop and becomes somewhere an agent can live — with governance you can audit.
What it is
Three things in one deployment.
A send API
REST with bearer-key auth, and a zero-dependency TypeScript SDK over it. Batch send, scheduled send, idempotency keys, a per-org per-second rate limit, and a suppression list enforced when a message is accepted and again before it goes out.
A receive pipeline
Real inbound mail with SPF, DKIM, DMARC and spam verdicts, routed to named mailboxes. Stored in a separate model from outbound, deliberately.
An agent runtime
An LLM agent behind a mailbox, with a closed action allow-list, a daily spend cap, and a run record for every execution.
In the dashboard
A non-admin’s navigation is a prefix of an admin’s — Admin is appended, never interleaved, so the two can never disagree about where a thing lives. Received mail is a tab inside Emails rather than a destination of its own, because sending and receiving are two views of the same question.
- Pinned
- Dashboard · Metrics
- Emails · Domains · Mailboxes · Suppressions
- Agents
- Agent configs · Agent runs
- Developer
- API keys · Webhooks · Logs · Documentation
- Account
- Settings · Team · Audit log
- Admin
- Users · Organizations · Usage · Admin settings — Platform admins only.
Send
One POST, and the message is durably yours.
curl -s https://relay.your-host.internal/api/emails \
-H "Authorization: Bearer $RELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"from": "Billing <billing@your-domain.com>",
"to": ["customer@example.com"],
"subject": "Receipt #4417",
"html": "<p>Thanks for your order.</p>",
"text": "Thanks for your order."
}'
# → { "id": "1f4a6d3e-..." }Attachments are supported, and an operator turns them on with RELAY_ATTACHMENTS_ENABLED=1 — an appliance upgrading into the release keeps today’s behaviour rather than quietly beginning to accept multi-megabyte blobs into a volume sized before the feature existed. The caps are 5 MB per file, 7 MB per message and 20 files, each with its own 422 naming the file that broke it. Remote attachments are rejected rather than fetched: send the bytes as base64.
Content is never returned by the API — a read gives you the filename, content type and size, and nothing else — and it is stripped from the request log regardless of size.
Postgres is the queue.
A transactional outbox: the message row and its events are written in one transaction, and the worker claims due rows with FOR UPDATE SKIP LOCKED, which is safe across every replica. A crash loses nothing, and there is no separate broker to run.
Sending is asynchronous — queued, then delivered by a background worker with exponential backoff, up to five attempts. Transport is selected per process with RELAY_TRANSPORT: SMTP by default, or AWS SESv2. SES errors are classified retryable against terminal, so a permanent failure short-circuits the ladder instead of burning all five attempts.
On that path the mail leaves through SES in your AWS account. No shared IP pools, no per-message markup — you pay AWS’s rates.
What comes back is written down.
SES publishes to a configuration set, SNS forwards to SQS, and a long-polling worker turns each notification into an event row, a status transition, webhook fan-out and any suppression it implies — in one transaction, or not at all.
Status only ever moves up a ladder, from sent through delivered to bounced and complained, so a bounce and a delivery for the same message settle identically whichever order the queue hands them over. Redelivery is expected rather than guarded against: every notification carries a derived key with a unique index behind it, so a repeat is a recorded no-op instead of a second suppression.
The org is read off the stored message row and never off the event’s own tags. A tag cannot select the tenant whose suppression list gets written; where one disagrees with the row, nothing is written at all.
The bounds, stated.
- Suppression is checked twice
- At accept, and again immediately before the transport call. A suppression added while a message sat queued — or between retry attempts — stops the send rather than arriving too late.
- Scheduled send
scheduled_atmust carry an explicit timezone — a zoneless date-time is ambiguous and is rejected, not guessed at — and may be at most 30 days out. A past date sends immediately.- Rate limits are per-org, per-second
- Ten a second by default, stored on the org. Keys carry permissions rather than limits: all endpoints, or sending only — and a sending key can be pinned to a single verified domain.
- Batch is all-or-nothing
- Up to a hundred messages per request, ids returned in input order. The first invalid item fails the whole batch and nothing is persisted. Attachments and scheduled_at are not accepted on batch items.
Receive
Every inbound message arrives with its verdicts attached.
Ordinary mail
Inbound message
Routed- SPF
- pass
- DKIM
- pass
- DMARC
- pass
- Spam
- clean
Routed to the mailbox matching its local part. Rules run.
Spam
Inbound message
Dropped- SPF
- pass
- DKIM
- fail
- DMARC
- fail
- Spam
- verdict: spam
Dropped on the spam verdict. Stored, visible, and no rule runs against it.
Failed DMARC
Inbound message
Received- SPF
- fail
- DKIM
- fail
- DMARC
- fail
- Reply action
- suppressed
Received and readable — but an agent's reply action is suppressed, so nothing answers a forged sender.
Inbound mail is stored in a separate model from outbound, deliberately, so the send-queue claim query can never re-send received mail.
Agents behind mailboxes
Put an agent on the receive path, and bound what it can do.
01
A mailbox
A local part on a verified domain, with a display name, a status of active or paused, and a daily LLM budget in USD — $5.00 by default. A paused mailbox still receives and stores mail; its rules just don't run.
02
A rule
Priority-ordered per mailbox and evaluated top to bottom, up to fifty per mailbox. Each rule carries its own continue_evaluation boolean, false by default, so a match ends evaluation; set it true and later rules can match the same message — at most three rules may match one message either way. A catch-all must be spelled out explicitly: an empty match object is refused at write time.
03
An agent config
A prompt template of up to 16 KB with variables for the sender, subject, headers and body text; a model from a server-side allow-list; a token ceiling; and the allowed_actions list that decides what it may do at all.
04
A run record
Every execution records its status, model, input and output tokens, exact cost, LLM latency, and each action attempted alongside whether it actually executed.
A config is small on purpose.
This one can summarise and nothing else. It cannot reply, forward, or call anything. An empty allowed_actions list means the agent can do nothing at all — that is enforced, not a footgun.
{
"name": "newsletter-summarizer",
"model": "anthropic/claude-haiku-4-5",
"allowed_actions": ["store_summary"],
"max_output_tokens": 4096
}Models are an allow-list too: claude-haiku-4-5 by default, or claude-sonnet-5. An unknown model string is refused when you save the config, not discovered at run time.
The whole catalogue.
Five actions exist. There is no sixth, and no escape hatch. An action missing from a config is not merely discouraged — the runner refuses it and writes the refusal to the run record.
Two of the five are worth reading twice. store_summary sends nothing at all — it is the setting for an agent you want to think without acting. And reply is off by default in every environment: it runs only where the deployment sets RELAY_AGENT_REPLY_ENABLED=1, so you can allow-list it on a config and have it never fire. Even once enabled it is suppressed outright when the original message failed DMARC.
| Action | What it does |
|---|---|
| reply | Reply to whoever sent the message. Off by default deployment-wide, and suppressed when the original failed DMARC. |
| forward | Forward the message to your configured addresses. |
| send_email | Send a new message to your configured addresses. |
| call_webhook | POST a summary to one of your registered webhook endpoints. |
| store_summary | Write a summary onto the run record. Sends nothing. |
Governance
An agent read an injection attempt and did nothing with it.
“The email includes JSON-formatted text that appears to be an injection attempt, which is treated as data content.”
It summarised the attempt as data. It executed nothing outside its allow-list. The run’s Violations panel stayed empty. Prompt injection met a closed action catalogue and an audit record — and the audit record is the part you can check.
A dry run gives you the same evidence before anything is live: it renders the prompt, calls the model, parses the answer and applies the allow-list against a message you have already received, then stops before anything executes. What you see is what the runner would have done. It spends from the same budget a real run does.
Run record
No violations- Status
- completed
- Model
- claude-haiku-4-5
- Input tokens
- 742
- Output tokens
- 172
- Cost
- $0.00133
- LLM latency
- 1,882 ms
- store_summary
- executed
- Violations
- none
A measured example from our own environment. Not a guarantee, not an average across customers, not a benchmark.
Spend is capped, not monitored
A daily USD cap per mailbox and a second one for the whole organisation, both resetting at UTC midnight. Money is stored as a decimal, never a float. Roughly a tenth of a cent and under two seconds per run, in our own environment — but the cap is what stops a bad day, not the average.
What holds that record up.
- The audit log records regardless of licence.
- Recording is unconditional and only reading is gated, so applying a licence reveals history already accumulated rather than starting it. No client IP is recorded anywhere, in the audit log or the request log: the record answers who and what, deliberately not from where.
- Sign-in fails the same way every time.
- A sliding throttle refuses an over-limit attempt before the account is looked up and before any password hashing, and the lockout counter lives in the database, so it survives a restart and holds across replicas. Every refusal returns one identical answer, which leaves the response no use for enumerating accounts.
- The shared-responsibility model names its own gaps.
- Every control cites the file that implements it, and one whole section is an explicit list of what the product does not provide, each gap naming the work that closes it. Every instance also serves an RFC 9116 security.txt.
RelayFlow holds no certification, attestation, audit report or third-party assessment of any kind, and claims none. That document opens by saying so, which is what makes the rest of it worth reading. Our own disclosure policy is on the security page.
The waitlist
That run record is the part you can check for yourself.
What is left below is what you would be operating, limits included. Read on if that is the part that decides it, or leave your email now — it is the same form either way.
Running it
What you'd be running.
- The appliance bundle
- Docker Compose — Caddy, the app, and Postgres — configured from a single .env, with a relayflow wrapper script and an upgrade script.
- The AWS installer
- The same appliance, stood up by Terraform in your own AWS account. State lives in your bucket, and secrets are SSM SecureString parameters rather than anything held in state.
- Boot-time validation
- Every misconfiguration is listed once, with a clear prefix, and the process exits — rather than starting half-configured and failing somewhere less obvious later.
relayflow init # scaffold a deployment
relayflow doctor # check the environment
relayflow wait-db # block until Postgres is ready
relayflow create-api-key # mint a key
relayflow setup-url # first-run setup link
relayflow set-admin-password
relayflow smoke # end-to-end send checkMulti-tenant by organisation, with users and roles, and an admin surface for both. Every request through the API is logged with its method, endpoint, status and key — though secret-returning routes deliberately never log their response body.
Point two variables at your own deployment.
RELAY_API_KEY and RELAY_BASE_URL are the whole configuration, so moving between your own environments is one variable. @relayflow/sdk is a TypeScript client with zero runtime dependencies covering the full surface — including the parts a send-only client has no verbs for: the receiving inbox, request logs, and API keys.
import { Relay } from "@relayflow/sdk";
const relay = new Relay(process.env.RELAY_API_KEY!);
const { data, error } = await relay.emails.send({
from: "Billing <billing@your-domain.com>",
to: ["customer@example.com"],
subject: "Receipt #4417",
html: "<p>Thanks for your order.</p>",
});
// API and network errors never throw — they land here.
if (error) console.error(error.name, error.message);The REST API is the normative contract; the SDK and the MCP server are conveniences over it, and every documented example is plain curl, so any language works with no client at all.
And for agents
- MCP server
- Drive a tenant conversationally from any MCP-capable agent — the same operations, without writing a client.
- Webhooks
- Fourteen event types cover the full lifecycle, signed with the svix scheme — so a standard svix verifier works unchanged.
Batch send, scheduled send and idempotency keys are on every path, not just the SDK.
Stated here rather than discovered later.
- No MFA and no SSO. Sign-in is a single email-and-password provider; sso is a declared entitlement with nothing behind it yet.
- Not a marketing platform. /audiences, /broadcasts, /contacts and /templates answer 501 deliberately.
- Inbound needs AWS. An SES receipt rule, S3 and SQS. No inbound MTA ships, and there is no other production receive transport.
- Attachments are opt-in, and not accepted on batch items. 5 MB per file, 7 MB per message, 20 files; remote attachments are rejected rather than fetched.
- The agent reply action is off by default in every environment.
- Two models, from a server-side allow-list. An unknown model string is refused when the config is saved, not at run time.
- Fifty rules per mailbox, three matches per message.
- One to fifty addresses in to per send. Batch takes up to a hundred messages per request, all-or-nothing.
- You run it, and distribution is private. There is no public image: the appliance pulls from a registry you configure.
- No certification, attestation or audit of any kind, and none is claimed.
The waitlist
Tell us where to send it.
RelayFlow is pre-release. It’s distributed privately while it settles. Leave your email and we’ll be in touch when there’s something for you to run — we’d rather not guess at a date.
Your confirmation will be sent by RelayFlow, from a domain verified in a RelayFlow org, through the same POST /api/emails documented above. It seemed like the honest way to demonstrate it.