Introducing Faira Memory

OEOvie Esiekpe · CPO & Co-founder

For most of my career I believed the hard part of customer software was the interface. Get the inbox right: every channel in one thread list, one place to reply, and everything else follows. For human teams, that's very nearly true. Someone looking at a unified inbox does the difficult work themselves, unconsciously: they see a WhatsApp message from an unfamiliar mobile number, remember Tuesday's call, recognise the name, and treat it as one customer with one problem. The interface didn't unify anything. The person did.

That assumption collapses the moment the thing reading the inbox isn't a person.

Over the last two years, conversational models have crossed a threshold. A voice agent can now hold a call that a customer doesn't resent. It can interrupt gracefully, handle someone talking over it, follow a change of subject. Once that stopped being the constraint, the constraint moved somewhere much less glamorous: retrieval. An agent that speaks beautifully and knows nothing is worse than a voicemail, because it wastes the customer's time before failing them.

Today we're introducing the layer that sits underneath every Faira agent. We call it Faira Memory.

MobileInstagramEmailBrowseriPadWhatsApp
One customer. Five channels. Nothing linking them.

The benchmark is a person, not a database

The standard we build against isn't a system of record. It's the receptionist who has worked somewhere for eleven years, who knows the voice before the name, remembers that this customer's last appointment was rescheduled twice, knows which of the two partners they prefer to see, and never once asks them to start from the beginning.

Nobody has ever described that person as a database, and the reason matters. What makes them good isn't storage. It's that recall arrives at conversational speed, attached to the right person, and shaped by everything the business has learned about this particular customer.

What Faira remembers

A booking system stores appointments. Memory holds the shape of a relationship.

The last four appointments and what happened to each one: kept, moved, missed, paid late. The practitioner they asked for the first time, and the one they've quietly asked for ever since. The treatment they enquired about in spring, were priced for, and didn't book. The figure they were quoted, which the agent has to honour whether or anyone wrote it down. That they reply to WhatsApp within minutes and never answer the phone before lunch. That they cancelled once, and why.

None of that is a field someone remembered to fill in. It's the residue of every conversation the business has ever had with that person, on every channel, resolved to one customer and available inside the turn.

Faira runs across telephone, WhatsApp, Instagram, web chat, our agent websites, and offline capture, with telephony coverage in 32 countries. Meeting that standard across all of it imposes three requirements that software designed to be read was never designed to meet.

Retrieval inside a breath. A screen can take two seconds to load and nobody complains. A person is reading, and reading absorbs latency. A receptionist cannot pause for two seconds mid-sentence while a customer is speaking. Silence on a phone call is not a loading state; it's a judgement about your business. Our budget for a full context retrieval is under 300 milliseconds, inside the turn, every turn. That single number rules out most of the standard architecture.

Identity with no shared key. The same customer arrives as a mobile number, an Instagram handle, an email address, a browser session, and a name typed onto an iPad at the front desk. Nothing links them. There is no foreign key. Most software resolves this by asking someone to merge duplicates later, fine for reporting, useless in a live conversation, because the moment to know who someone is has already passed by the time anyone opens a screen.

Concurrent, cross-channel writes. A customer messages on WhatsApp while their partner calls the same practice about the same appointment. Two Faira agents, two channels, one underlying customer, both writing at once. In software built around human editing speed this is rare enough to ignore. In ours it happens constantly, and getting it wrong means one agent confidently contradicting another.

Resolution as a conversational strategy

The most interesting part of Memory isn't the store. It's what we do with uncertainty.

Faira resolves identity probabilistically and continuously, scoring every signal it has: number, handle, name, service interest, timing, device, prior conversation shape, into a confidence that this contact is someone we already know. Crucially, that confidence isn't a hidden property. It changes how the agent talks.

Above a high threshold, Faira simply uses what it knows. It doesn't ask a returning patient to explain themselves again; it picks up where the last conversation stopped.

In the middle band, Faira verifies conversationally rather than guessing silently. "Am I right that you called last week about the Invisalign consultation?" A question a good receptionist asks, which happens also to be a confirmation event that resolves the identity graph. When the customer says yes, the merge is made and it's defensible.

Below the threshold, Faira starts clean and says nothing. The failure mode we care most about is not forgetting someone. It's confidently addressing a stranger as someone else, and that's the failure no amount of speed recovers from.

Why we didn't bolt on a vector store

The conventional way to add semantic recall is to replicate data into a separate vector store and query it alongside everything else. It works, and for asynchronous workloads it's entirely reasonable.

It doesn't survive a live call. Replication lags. On a screen, a few seconds of lag is invisible. In a conversation, it means the agent answering the phone doesn't know about the WhatsApp message the same customer sent ninety seconds ago, and the customer absolutely does.

Split stores also force the agent to reconcile structured lookups against unstructured recall at request time, inside the same 300 milliseconds it has already spent. Memory keeps structured attributes, conversation history and semantic embeddings in one consistent path, written transactionally. When Faira answers a call, what it knows is what happened, including what happened while the phone was ringing.

Reversible, and auditable

Recall that can't be corrected isn't an asset. Every merge Memory makes is reversible. Every merge records the evidence that produced it. No merge is ever made on a signal a person couldn't audit afterwards, and every recalled detail can be traced back to the conversation it came from.

The record belongs to the business, not to us, and not to whichever member of staff happened to take the call.

Why the order we built in mattered

The straightforward way to ship in this space is agent-first: take a model, point it at a channel, launch. The identity problem doesn't appear on day one. It appears the day a second channel is added, and by then the data model is set, and the fix is a migration nobody has time for.

We built Memory first and the agents on top of it. That was slower, and it's the reason Faira behaves as one front desk rather than five products sharing a logo.

What's next

Memory sits underneath every Faira agent, on every channel we support. The roadmap extends it in two directions: outward, so the systems your business already runs on can read and write the same customer context Faira does; and deeper, so that the resolution improves with every confirmed conversation rather than every configuration change.

Recall is only half of it. What a long-tenured receptionist also has is judgement about what to say next, learned over years of watching which conversations turned into work. That's Faira Instinct, and it's built on this.

The front desk has always been where a business's memory really lives, in the person who recognises a returning customer before they've finished saying hello.

We've spent two years building software that can do the same.

Stay on top of revenue opportunity

Go live in 10 minutes. No technical setup required.

Start free trial