An inbox agent that drafts but never sends
Most of what lands in an inbox needs nothing from you, and the few emails that do arrive mixed in with the rest. I stopped checking mail and let a script read it instead: every new email goes through a small LLM, gets one of five buckets, and the buckets decide what reaches my phone. Urgent items are pushed within ten minutes. Everything else waits for a digest at seven in the morning.
whalemail is that script, open source under Apache 2.0. It is a little over two thousand lines of Python including tests, with two dependencies, both for OAuth. This post is about the three decisions that shaped it.
Five buckets, ordered
The classifier has five outputs: action, decision, fyi, verify, noise. The names are the delivery policy. Action means a service stops or money is lost if you do nothing, and it is pushed as it arrives. Decision has a deadline but tolerates a morning’s delay. Fyi is settled and gets one line in the digest. Noise is counted and never listed.
Verify is the bucket that took longest to get right. It holds anything about money, accounts or logins that reads like phishing, and it exists because “your payout method needs updating” is the shape of both a real Stripe email and a fake one. The classifier flags the shape, and the digest line carries a fixed note: check through the official app or site, do not follow the link. The rules tell the model never to suggest a link, and the note is added by the code regardless of what the model wrote.
When several buckets apply, the order is verify > action > decision > fyi > noise. Suspicious beats urgent, so a phishing email dressed as a payment failure is never pushed as an alert.
Over-report on purpose
An email classifier fails in two directions. It can push things you did not need to see, which costs a glance. It can drop something you needed, which you find out about later, from the consequence. The asymmetry decides the defaults.
If the model returns a bucket that is not one of the five, the email becomes fyi. If the model skips an index in its JSON, that email becomes fyi. The rules file, which is the system prompt, says the same thing in words: anything about money, payouts or account status moves one bucket up when unsure, never down to noise. In practice this means the digest sometimes lists a receipt as fyi that could have been noise. I have not once found something important in the noise count.
The rules file is split in two. rules.md holds the bucket definitions and the safety lines and ships with the repository. rules.local.md is git-ignored and appended at run time; mine says which of my own Gmail labels mean “unimportant” and which banks I actually use. The classifier gets both. The repository gets one.
Keep the send button human
The same Telegram group has an agent. You write “find the last invoice from Cloudflare and draft a reply asking for the VAT number”, it searches, reads the thread, and calls draft_reply. That tool writes to the Gmail Drafts folder and nothing else. The bot then posts a card with the recipient, the subject and the body, and two buttons: send and cancel. Only tapping send calls drafts.send. Cancel leaves the draft in Gmail for editing by hand.
An agent that can send email on its own has a failure mode with a reply-all in it, and I would rather tap a button forty times than explain one of those. The tool schema the model sees has no send function at all, so there is no prompt that reaches it.
The OAuth scopes follow the same line. whalemail asks for gmail.modify and gmail.settings.basic, which cover reading, labels, drafts and filters. It does not ask for https://mail.google.com/, the full-access scope. The difference is permanent deletion. A leaked token from this project can read mail and create drafts; it cannot empty an account.
Small things that turned out to matter
The two-hour heartbeat deletes its previous message before posting a new one, so the group never fills with counts. It is delivered silently and exists so that “no alert” and “the job died” look different.
Timestamps in the digest show only the time for today’s mail and the date for anything older, because a digest at seven covers the night and “23:40” alone is ambiguous.
The fetch query defaults to everything received, which includes mail that Gmail filters archive on arrival. Those filters tend to catch exactly the automated mail this tool is for.
Messages and model output follow one setting, WHALEMAIL_LANGUAGE, so the same code runs my Chinese digest and an English one.
Running it
You need a Google Cloud OAuth client, any endpoint that speaks /chat/completions, and a Telegram bot. The default model is a cheap one on OpenRouter; classification does not need more. On macOS a script installs launchd jobs for the poll, the heartbeat, the digest and the bot, with the schedule read from .env. On Linux the same six commands map onto cron.
The repository is at github.com/openwhale-labs/whalemail.