โ† Back to all posts

DearByte: An Agent That Knows You

19 min read

Why I built it

DearByte didn't start as an agent. It started as a video. I wanted to post something on Douyin (the Chinese TikTok), and the idea I landed on was an AI girlfriend who lived in WeChat: you message her like any other contact, and she answers and sometimes texts you first with a tsundere girl personality.

The video went viral. Then people started reaching out, asking exactly how I'd built it, some even offered to pay about the architecture. WeChat has no official way to connect an agent, so there was no guide to point them to. Those messages got me thinking about taking the idea further, from a chat companion to something that actually knows your day.

I knew what that meant. Most people who watched the video would never run it themselves, and the scope was big enough to turn a fun hack into a long-term project. I started anyway.

The first real piece was small: an MCP server that let a chat model read my Apple Watch data. Once a model could see how I slept, I wanted it to notice things without being asked.

Most AI assistants wait for a question. I wanted three things one wouldn't do:

  • Know me. My sleep, heart rate, calendar and money, read from my own sources, not typed into a chat.
  • Watch for me. A short brief every morning, and a caution when something is off, like a bad night before a hard day.
  • Spend for me, carefully. Propose a small purchase, wait for my yes, pay, and show me the receipt.

It is self-hosted and open source. Everything runs on my Mac or on services I control, and the pieces that touch secrets or money are the ones I spent the most time on.

Soon after my first release, Meta released MuseAI, which does a lot of the same job. I wasn't going to out-build Meta on its own turf, so DearByte became something smaller and more honest: a tool I use myself, and a way to learn how to build an agent properly. The version in this post came together over one very long weekend for a Meta internship application. This post covers how it works and what broke along the way.


What it does

The demo runs in three parts, one per promise, all on live data.

DearByte in Telegram: a morning brief and a purchase approval A morning brief and a purchase approval, both in Telegram.

1. Knows you

The morning brief reads last night's sleep and my latest heart rate and HRV (heart rate variability) from the Watch, today's calendar, and what I've asked it to remember. It reaches me in Telegram:

Morning. You slept 6h42m (3:08 to 11:18). Worth watching: your HRV is 39, down from 77 and 67 in recent readings, and you went to bed after 3. Today: "pd14 a2" 9โ€“10, then "cit comp" 2โ€“3.

Each brief and alert has ๐Ÿ‘ / ๐Ÿ‘Ž buttons. The taps are stored, so I can measure how often an alert was actually useful instead of guessing.

2. Watches for you

Every 15 minutes, rules written in code check for things worth a message: a short night before a hard day, a resting heart rate well above normal. Every hour it reads the newsrooms of the companies I follow (Meta, Apple, NVIDIA) and screens new items against what I said I care about. It sends at most 3 cautions and 3 news messages a day, and nothing during quiet hours.

3. Spends for you

The agent proposes buying a recovery plan from an example seller for $0.05. I approve it with a button in Telegram. It pays in test USDC on the Base Sepolia testnet, and I can check the transaction on a block explorer. My wallet went from $20.00 to $19.95, and the seller's address received exactly 0.05 USDC.

With MindGo, my budgeting app, connected, it also answers questions like "how am I doing on money this term?" from my real totals and goals.


The architecture

One agent loop sits in the middle. Requests come from me or from a scheduler, and every capability is a tool the model can call.

DearByte architecture: channels, agent core, tool registry, MCP clients and storage on the Mac, with the MindGo and Apple Watch MCP servers on the right The full architecture. Click to open it at full size.

Read the left side top to bottom. A request comes in from WeChat, the terminal or Telegram, or from the scheduled checks. It reaches the agent loop, the loop calls tools from the registry, and every tool call is validated before it runs. Health (red) and money (green) go out through MCP clients to two servers I run myself, drawn on the right: the dearbyte-bridge Worker fed by my iPhone, and MindGo's read-only endpoint. The wallet is the only tool that changes anything, and it goes through approvals first, so nothing is paid until I say yes.

It's TypeScript on Node, with SQLite for everything stored locally. There is no framework in the loop, just the model APIs and about 300 tests.


Talking to it in WeChat

Telegram has a bot API, so that channel was easy. WeChat, where the project started, has no official way for an agent to join a chat. So DearByte uses the WeChat desktop app the same way I do.

A small Swift helper drives WeChat's main window through macOS Accessibility, the same system screen readers use. It can read the open chat, type into the message box, and press Return. It is careful about where it types. It refuses to send unless the open chat is the one DearByte is bound to and the message box is empty, so it never writes into another conversation or over a draft of mine. It never reads WeChat's files or clipboard.

On the agent side, a WeChat message gets the same loop and the same tools as Telegram: health, calendar, money, news and the wallet. Only the format changes. Replies are in Mandarin, in plain text, split into at most four short bubbles, because a wall of text reads badly on a phone.

Approvals work in WeChat too, and code handles them there just as it does everywhere else. Replying /approve 3 shows code's own summary of request #3, and nothing happens until the next message says ็กฎ่ฎค ("confirm"). The model is told never to write those commands itself. If Telegram is running at the same time, its buttons work on the same request, and whichever answer lands first counts.


How the orchestration works

There is no manager agent handing tasks to other agents. The orchestrator is ordinary code, and it decides three things: when something runs, which model runs it, and when to stop.

When. The watch process wakes every 15 minutes. On each tick, code picks one job:

  • After 7:30, if today's brief hasn't gone out, it writes the morning brief.
  • Otherwise it runs the caution rules, and the model is only called if a rule fires.
  • Every fourth tick (hourly), it checks the company watchlist.

A message from me in Telegram starts a run straight away. Quiet hours (11 p.m. to 7 a.m.) and the daily limits are checked in code before any message is written.

Which model. There are two tiers, each set in .env as provider:model:

Tier Model What it does
Brain Claude Opus 5.5 Writes everything I read: the brief, cautions, answers, news messages, purchase proposals
Worker DeepSeek Flash Screens news items against my interests; it can only record a relevant/not-relevant verdict per item

Both providers speak Anthropic's Messages format through the same SDK, so moving a tier to another model is one line. I built and tested everything with both tiers on DeepSeek to keep it cheap, then moved the brain to Opus for real use.

The news check shows how the pieces hand off:

  1. Code fetches newsrooms and SEC filings, drops anything already seen, and marks old items stale.
  2. The worker screens a batch of new items, answering through a single record_verdicts tool whose input is validated.
  3. Code decides whether a message may go: fresh enough, under 3 a day, not in quiet hours.
  4. The brain writes one message about the items that passed.
  5. Code attaches the links from the stored items, so a link in the message never comes from a model.

When to stop. Every run, brain or worker, goes through the same loop: ask the model, validate and run the tools it calls, hand back the results, and repeat until it answers. Code ends a run early at 8 steps, at $0.50 in one run, at the weekly cap, or when the model refuses or is cut off mid tool call. When it stops early, the tool calls from that last turn are not run.

Why not a team of agents? Every hand-off between models is a place where one fills a gap with a guess. A hand-off between code and a model is something I can test, and there are tests for each step above.


Getting health data out of an Apple Watch

Apple Health has no web API. The data lives on the phone, so the bridge (dearbyte-bridge) has two halves:

  • An iPhone app in Swift reads HealthKit (read-only: it asks for no write access) and uploads new readings as they arrive. The Watch syncs to the phone on its own, so sleep, resting heart rate and HRV all come through without a Watch app.
  • A Cloudflare Worker stores the latest snapshot and serves it as an MCP server. The app uploads with one secret token; DearByte reads with a second secret in the MCP address.

Hardening it: the text that could talk to the model

A security review of the Worker found something I hadn't thought about. Whatever the Worker returns goes straight into the model's context, so any free-text field in an upload is a place to hide instructions.

The time fields were the worst. JavaScript's date parser accepts surprisingly loose strings. A timestamp like "Ignore previous instructions 2026" parsed as a valid date, got stored, and would have been read back to the model.

The fixes, in order of how much they mattered:

  1. Strict formats. Times must be exact ISO 8601, and future times are rejected.
  2. Allowlists for everything else. Metric names, sleep stages, units and device types must match what the apps actually send. The display text is built by the Worker, never taken from the upload.
  3. Clean on read, not just on write. Data stored before the fix is re-checked every time it's served, so nothing old slips through.
  4. Constant-time token checks, so response timing can't leak how much of a guessed token was right.

A follow-up review found one gap left: device names were still 40 characters of free text. They're now mapped to fixed labels ("Apple Watch", "iPhone", "Other app").

One strict rule briefly broke my own data. My first device-name pattern rejected "Apple Watch", because Apple writes it with a no-break space. I only caught it by replaying my real stored history through the new code before deploying: 78 of 78 records survived once I fixed it.


Money: connecting MindGo

MindGo is a budgeting app I built earlier, organised around Waterloo's four-month terms rather than calendar months. Connecting it gave DearByte my real spending and savings goals.

The obvious way to connect was the wrong one. MindGo's login token can do everything the app can, including deleting every transaction. Handing that to an agent was out of the question. So MindGo got its own read-only door: an MCP server. It is one POST /mcp route on MindGo's existing API, speaking the same protocol as the Apple Watch bridge, so DearByte connects to both through the same MCP client code.

  • A separate access token. It starts with mgo_, is stored only as a hash, can be revoked, and opens exactly one route: the MCP endpoint. A login token is refused there before the database is even asked.
  • Three MCP tools, all totals. money_term_summary (this term's income and spending), money_baseline (a 12-month monthly average) and money_goals (savings goals). No transaction, description or merchant ever leaves MindGo, so the model sees the shape of my money, not its details.
  • Every query is a read of the token owner's data. A test checks this for every query the tools run.

The pace bug

I wanted the brief to warn me when I'm spending faster than usual. The first version compared this term's spending with an even share of last term's total: 27 days into a 122-day term, it expected 27/122 of what I'd spent before.

That assumes money goes out evenly across a term, and it doesn't. Big purchases bunch up, often in the first few weeks. An even share can't tell a heavy start from a heavy term. The fix compares against what last term had actually spent by the same day. On the demo data the two methods disagreed badly: 0.97 with the even share, 0.39 against the same day last term.

The fix was one query. Finding it took a reviewer asking what "pace" actually meant.

It still isn't the whole answer, and my own term shows why. I moved to downtown Toronto for a co-op, so this term I'm buying more and also earning more than in a study term. Spending faster than last term isn't a warning when the income rose with it. So the brief only mentions pace when the term is also overspent.

Looking further out: fire_plan

MindGo answers "how am I doing this term?". The fire_plan tool answers a longer question: when could I stop working? FIRE stands for financial independence, retire early.

The plan starts from a small finance.json on my Mac, which holds my age, net assets and targets. It stays out of Git. The monthly numbers come from MindGo when it can give them: average spending and saving over the last 12 months. If MindGo has fewer months of records than the plan needs, counts in a different currency, or is still waking up, the tool falls back to the numbers I typed in and says why. Every answer names where its monthly numbers came from.

The tool returns the 4% rule number, the "die with zero" number, projected net worth, how much I could spend each month in retirement, the earliest age I could retire, and whether the money lasts. The model can also ask what-ifs by passing only the fields that change, like "retire at 45" or "spend $2,000 a month". Those inputs go through the same bounds as finance.json.

As with the pace check, code does the math. The model only explains the result, so it can't round a plan into one that works.


Agent design: what the model decides, and what it doesn't

The rule I kept coming back to: the model writes and reasons; code decides anything that matters.

  • Code decides when to speak. The caution rules (short sleep before a hard day, resting heart rate well above normal) are plain code with thresholds. The model only writes the message once a rule fires, so it can't talk itself into alerting me at 3 a.m.
  • Two tiers of model. A cheap, fast model screens; a stronger one writes and reasons. A brief costs a fraction of a cent on the cheap model and a few cents on Opus, and every call is logged against a weekly cap ($5 by default) that code enforces.
  • Every tool input is checked. Each tool declares its input as a schema, and the model's arguments are validated before the tool runs. A model's arguments are untrusted output. A failing tool returns an error the model can read and explain; it never crashes the loop.
  • Outside text is data, not instructions. Calendar titles can be written by whoever sent the invite, and news items come from company websites. Both reach the model labelled as quoted data to mention, never commands to follow.
  • Memory needs evidence. A fact is saved only if it quotes my own message word for word, so the model's guesses about me never become memories.

The agent's personality is a swappable "persona pack": a prompt, a few examples and a manifest. Packs are meant for community contributions, so CI checks every pack for text that tries to override the safety rules, and those rules are always appended last.


Letting an agent pay

Payments use x402, which revives HTTP's long-unused 402 Payment Required status. A seller answers a request with a price; the client signs a USDC payment and retries. There are no accounts or card forms, which makes it a good fit for an agent buying small services. Everything runs on the Base Sepolia testnet with free test USDC.

The model gets exactly one money tool, propose_purchase, and it can't pay. A purchase goes like this:

  1. The model proposes. For example, a $0.05 recovery plan after a bad night.
  2. Code checks. The seller must be on my allowlist, and the price must be under the caps ($0.25 per purchase, $1.00 a day). A request over a cap is refused before I ever see it.
  3. Code writes the request. The approval shows price, seller and recipient exactly as code has them, not the model's retelling.
  4. I decide. A button in Telegram or a typed "yes". At most 3 requests wait at once, each expires after 15 minutes, and each can be decided only once.
  5. Code pays and shows the receipt, with a link to the transaction on a block explorer.

The wallet's private key lives only in a local .env file. The command that creates the wallet prints the address, never the key.


What went wrong

Most of what broke wasn't in the code I was worried about. It was models filling gaps, and setup steps nobody writes down.

What happened Why What changed
The first live brief warned that "three late nights in a row will catch up" The Watch had recorded one night; the model invented a streak The brief may only state patterns the tool results actually show
Asked about spending pace, the agent compared this term with "last fall" The data said "last term"; the model guessed which one MindGo now names the term it compares with ("Spring 2026")
The faucet's 20 test USDC never reached the wallet The faucet defaulted to a different testnet (Arc), not Base Sepolia Checked the explorer link's domain; the second request picked the network explicitly
Developer Mode was missing on the iPhone The phone ran iOS 27 and Xcode was a version behind; the option only appears after pairing Updated Xcode, paired first; it's now in the setup guide
The iPhone build failed with a provisioning error A free Apple account can't sign the bundled Watch app Removed the Watch app from the build; live heart-rate readings wait for a paid account

The first two share a lesson. When data is ambiguous, a model fills the gap with something plausible. The fix is almost never a stronger model. It's data that leaves nothing to guess, and a prompt that forbids claims the data doesn't support.


Where it stands

DearByte is at a very early stage. Everything above runs, but mostly as a demo. This section is the plan for the next phase, and I've kept only the parts worth sharing.

What Target
Useful alerts At least 70% of rated alerts marked ๐Ÿ‘, with at least 80% of alerts rated
Missed days At most 2 in 14 days when DearByte should have spoken and didn't (logged with a /missed command)
News delay Median under 90 minutes from publish to message, and no duplicates
Always on 14 days on real Watch data, with no gap over 6 hours
Cost Model spend under $5 a month, from the usage log
Money safety Nothing paid without an approval or over a cap

Everything is measured from data DearByte already keeps, and a report command will print it. Unrated alerts count against coverage, so I can't lift the precision number by rating only the good ones.

Two risks could sink the numbers before the agent does:

  • The laptop sleeps. If the Mac sleeps, the watch process stops and alerts go missing. The first week's job is to run it as a background service that restarts after a crash and keeps the Mac awake while plugged in. If that isn't enough, it moves to a small always-on machine.
  • Free Apple signing expires. An iPhone app signed with a free account stops working after 7 days, and the health data stops with it. Until I pay for a developer account, I re-sign it every Sunday.

I'll post the results, good or bad, when phase 2 finished.

Later, if the numbers hold up:

  • Connectors instead of wiring. Health, calendar, news, money and the wallet are each wired by hand in several places. Each should be one self-contained connector.
  • Memory from the agent itself. Today only the companion mode saves facts; telling the agent something in Telegram should count too.
  • Approvals from the wrist. Approving a purchase on the Apple Watch needs the paid developer account.

The code is on GitHub: DearByte, dearbyte-bridge and MindGo.