OpenNous
All updates
ENGINEERING· August 5, 2026

How we built the Context Graph underneath the Revenue Layer

How we unify a revenue team's scattered communication into one context graph, and the engineering behind ingestion, identity resolution, and governance.

BGBennet Glinder

Nous unifies the communication a revenue team generates across every tool it uses and turns it into a working understanding of each account that both the team and its deployed agents can act on. The belief behind it is that the model an agent runs on is close to commoditized, and what determines whether it acts correctly is the quality of the context it is given. A context graph is the infrastructure that produces that context. It reconciles who is who across a dozen systems, keeps the record current as new activity arrives, and makes the state of any account available in a single request. This is how that works underneath, and how we built it.

The problem is that the data is fragmented

A sales rep is preparing for a follow-up call with a prospect. To walk in useful, they need to know what was discussed in the recent meetings, which emails went out and which ones went quiet, and who else on their team has engaged the account.

So they open four tools. Google Calendar to find when the meetings happened. Fireflies to pull the call recordings and transcripts. Salesforce to read the logged activities and notes. Gmail to scan the email threads. After fifteen minutes of moving between tabs and searching each system on its own, they have assembled a partial picture, and they still miss the parts that matter. The meeting a colleague ran last week never surfaces. The thread that went cold right after a specific call sits three screens deep.

This is more than an inefficiency. A revenue team holds an enormous volume of communication, and it lives inside separate systems with no single place to reason over all of it at once. When a deal stalls, the rep cannot readily answer what happened after the demo, or when engagement actually dropped off. The data exists. It is scattered across tools that were never designed to agree with each other. That same fragmentation makes the harder questions close to impossible, such as reading pipeline across the whole team and understanding why deals are not closing, because the evidence for that answer lives in five places at once.

The challenges that come with it

Four problems sit beneath that fifteen minutes, and each one has to be solved before an agent can be trusted with an account.

  1. Data lives in silos: every tool holds one slice of the relationship and none holds the complete state. The calendar records that a meeting occurred, the notetaker holds what was said, the CRM keeps what a person chose to log, and the inbox holds what was actually sent. No single system can answer a question that spans them.

  2. Every vendor uses a different schema: a meeting in the calendar, a recording in Fireflies, and an activity in Salesforce describe the same event under different field names, different identifiers, and different assumptions about time. Reconciling them by hand is where most internal data projects quietly stall.

  3. The same person appears under many identities: one prospect is a calendar attendee under a work email, a Fireflies participant under a personal address, a Salesforce contact under a CRM ID, and a LinkedIn profile under an opaque member URN. Treated naively that is four people, and treated correctly it is one. Getting that resolution right is the load-bearing problem of the entire system.

  4. The CRM is stale by design: it stores the fields a person remembered to fill in. It does not capture the objection raised on a call, the competitor named in an email, or the fact that a champion went quiet two weeks before the deal died. The record looks complete while missing precisely the context that would have changed the next move.

How we ingest data

Pollers and webhooks feed every connected source into one append-only observations table, and every derived layer is computed from it.

We take the raw activity from every connected tool and turn it into structured intel at two levels, the person and the account. The first principle we committed to is that we store the evidence rather than overwrite a field. Every inbound event, an email, a calendar invite, a call transcript, a reply on an outbound sequence, becomes an append-only row in one observations table, and we retain the original provider payload alongside it in a raw column. A database trigger rejects any delete and any edit to an observation once it lands, so the record of what was actually seen never changes underneath us. The current value of anything is always derived from that evidence, which means we can replay the full history and rebuild the graph whenever we need to.

Data arrives two ways. Pollers run every hour for the sources that expose a mailbox or a calendar, covering gmail, calendar, slack, and raw smtp. Webhook handlers take the sources that push events to us, and today that set spans the notetakers Fireflies and Fathom, the outbound platforms Instantly, Smartlead, lemlist, EmailBison and HeyReach, LinkedIn through Unipile, the schedulers Calendly and Cal.com, the visitor-identification feed RB2B, and Stripe. Each handler carries its own cleanup rules. Fireflies arrives as an array of sentences that we stitch into a clean speaker-labelled transcript and dedupe the participant list by lowercased email. Timestamps from every source are parsed to UTC at the edge before they are written, so the timeline stays coherent regardless of which tool a fact came in through.

The result is a single evidence spine in Postgres. Everything downstream, the identities, the claims, the scores, and the context an agent reads, is computed from it.

How we resolve one person across every tool

Identity resolution is what makes everything above it possible, and it is where we invested the most engineering. Every identifier we encounter, an email, a canonicalized LinkedIn URL, a CRM ID from HubSpot or Salesforce or Pipedrive, resolves to exactly one active entity in a workspace, an invariant enforced by a unique index in the database. When a new observation arrives, the resolver runs a three-step waterfall.

  1. Match on a stable identifier: when the incoming email, LinkedIn URL, or CRM ID already points at an entity, that is the person, and any new identifiers on the event are attached to them. This is the high-confidence path, and it handles the majority of traffic.

  2. Fall back to a corroborated name: a name on its own is never sufficient, because two people share a first name and a guess corrupts the graph. We anchor on the exact first name and then require a second signal, most often that the sender's work domain matches an identifier we already trust for that person. A match on a free mailbox such as gmail.com never corroborates, because it establishes nothing about who they are.

  3. Otherwise, create a new entity: when nothing corroborates, we would rather hold two records for a moment and reconcile them later than fuse two real people into one.

Four identifiers resolve through the three-step waterfall into one entity, with co-attendance and a nightly dedup as safety nets.

The hardest case is where a calendar event, a call recording, and a CRM activity all describe the same meeting under different identifiers. We resolve it through co-attendance. A recording is matched to an existing calendar booking when it falls within two hours of the same start time and the attendees align, guarded by a check that a contact's name actually appears in the meeting, so two back-to-back calls with generic titles never collapse into one. When duplicate records for the same person do slip through, a nightly sweep merges them, but only when they share a hard identifier such as the same email or LinkedIn URL. Name-only lookalikes are surfaced for a person to confirm and are never merged automatically. Every merge is lossless and reversible, so a mistake is always recoverable.

What we extract, and why it lets us find patterns

A contact's own words run through Haiku extraction and a semantic dedup into a claim that carries its confidence, epistemic class, and freshness.

Once a person is resolved, we read what they actually said. We run extraction only over the contact's own words, their replies, their messages, and the meetings they attended, and we deliberately exclude our team's outbound so we are learning from the customer rather than from ourselves. A Haiku model pulls a small number of durable facts from each interaction, capped at two per message and four per meeting so the signal stays high, and every fact is filed into a controlled taxonomy at the person or the account level, a goal, a pain, an objection, a competitor, an authority or budget or timeline signal, a stated preference, a relationship.

We focus on those durable, semantic facts deliberately, because they are what let us find patterns across accounts. Each fact is stored as a claim with an embedding, so before we add a new one we run a semantic search against what we already know and let the model decide whether this is genuinely new, an update to an existing fact, or a duplicate to skip. That same embedding layer is what turns which accounts raised the same objection this quarter into a query rather than a research project. Every claim also carries its own epistemics, a confidence, an epistemic class that records whether we observed it or inferred it, and a freshness that decays over time, so a fact that has gone stale announces itself rather than quietly misleading an agent.

What this looks like at scale

For a handful of accounts, you could nearly hold this in your head. A mid-sized revenue team cannot. Consider a fifteen-person team working a few hundred active accounts. Each account accumulates dozens of emails, several recorded calls, a steady stream of outbound touches across email and LinkedIn, and a trickle of CRM updates every week. That is tens of thousands of interactions a month, every one landing under a slightly different identity, in a slightly different schema, on a slightly different clock.

Without a context graph, answering what the state of an account is means the fifteen-minute reconstruction, repeated per rep, per account, indefinitely, and the cross-account questions stay unanswerable because no person is going to reconcile ten thousand records by hand. With a context graph, that same volume flows into one evidence spine, resolves to the right people and companies as it lands, and derives into a set of claims that a person or an agent reads in a single call. The work of joining, deduping, and dating the data happens continuously in the background instead of in a rep's head the moment before a call.

We keep pace with that volume by deriving incrementally rather than rebuilding in bulk. Every new observation enqueues a small job that recomputes only the claims it touches, drained roughly once a minute, and embeddings are backfilled a couple of minutes behind that. Heavier cross-account rollups run nightly. Because the claim layer is a regenerable cache over the immutable observations, we can always discard it and replay the evidence to rebuild the entire belief layer from scratch, which is the honest version of a rebuild that depends on no elaborate diffing.

What the context graph does for an agent

The point of all of this is the last mile, the moment an agent asks for context. When it does, we do not hand back raw rows and ask the model to reason over the mess. We assemble an intent-shaped record. A request for meeting_prep returns a different shape than one for draft_email or follow_up, each tuned for how far back to look, whether to include the buying group, and how large a token budget to spend. What comes back is the resolved entity, the ranked claims with their confidence and freshness attached, a tiered timeline, the stakeholders and the buying committee off the relationship graph, and the atomic notes we hold on the person. The account is assembled from the graph we already joined and scored, rather than reconstructed from five tools at query time, and that distinction is the entire point.

How companies actually use it, and why governance is the hard part

The moment you bring an agent into a revenue workflow, it stops being a passive tool. It reads data, surfaces insight, and takes action on behalf of a specific person. That raises a question that has nothing to do with model quality and everything to do with trust. When an agent acts for a rep, does it see exactly what that rep is permitted to see, and can it be used to reach what they are not.

The concrete version is uncomfortable, and it is the right test. One rep should not be able to read the CEO's emails through an agent. One salesperson should not see the quota or the deal notes on another rep's account. A payroll email in one person's inbox should never surface in a summary another employee requests. If the context is powerful, it is also dangerous, and it is only trustworthy when every read is scoped to what the person behind it is authorized to see.

Every request carries a viewer, and a single rawVisible check scopes what it can see, so an agent inherits exactly the visibility of the person it acts for.

So we built that scoping into the layer everything reads through. Each request carries a viewer, and a single function decides what raw content that viewer is allowed to see. An owner or admin sees everything in the workspace. A regular member sees the shared layer, the meetings, the derived facts, the signals, and the notes the team works from together, while the raw private bodies of another rep's emails and LinkedIn messages are redacted. The header and the timing stay visible, so a teammate still knows the touch happened, and the private content stays closed. The default fails closed, so when we cannot positively resolve who is asking, we under-share rather than leak. Because this lives in the same path the agents read through, an API key acts as the member who created it, which means an agent inherits exactly the visibility of the person it works for, and no more.

For the unstructured stores, the emails and the transcripts, ingestion is gated as well. Internal people are marked as internal so a teammate is never mined as if they were a prospect, and whole classes of ingestion are gated by plan before they ever run. We are honest about the shape of this. The isolation is enforced in the application at every read rather than delegated to the database, and the scoping happens as redaction when the data is read rather than as a role filter before it is stored. That is a deliberate design we can point at in the code, and it is the part that makes a revenue team comfortable directing an agent at their most sensitive workflows.

Every revenue team is quietly asking whether AI can be trusted with the functions that matter most, the account research, the outbound, the forecasting. The teams that arrive there will not be the ones with the most data. They will be the ones that invested in a shared understanding of what is true, and the controls to decide who is allowed to see it.

The results, and what it changes

By unifying a team's communication into one context graph, we let an agent answer questions in a single call that previously required opening four tools and reasoning over fragments by hand. Questions such as:

  • What was discussed in our last three conversations with this account, and what did they push back on?

  • Which emails went out after the demo, and where did the thread go cold?

  • Who on our team has engaged this account, and when did engagement actually drop off?

  • Which accounts across the pipeline raised the same objection this quarter?

Five tabs and fifteen minutes of manual reconstruction collapse into one get_context call that returns an assembled, ranked, dated record.

We answer those semantically, over resolved and scored claims, instead of retrieving raw rows and re-reasoning over the fragmentation on every request. For the person doing the work, that is the difference between a fifteen-minute reconstruction before every call and a record that is already assembled, ranked, and dated when they open it. Complete visibility into every touch on an account, in one place, at the person level and the account level, with the objections and competitors and stalls surfaced instead of buried.

This is what changes for a team. The context stops living in a rep's short-term memory and starts living in an infrastructure that keeps itself current, resolves who is who as data lands, and can be trusted to show each person exactly what they are permitted to see. That is the infrastructure AI-native teams need when they begin deploying agents at scale.

Recap

  • We store evidence and derive belief. Every event lands as an immutable observation, and the current state of an account is always computed from that record, so the graph can be replayed and rebuilt.

  • Identity resolution is the load-bearing problem. A three-step waterfall and co-attendance matching turn four fragmented identities into one person without ever fusing two real people by accident.

  • Governance is what makes the context usable. Every read is scoped to the viewer, agents inherit the visibility of the person they act for, and the default fails closed.

If you want to see this run on your own data, that is what we built Nous to do. Point it at the tools your revenue team already uses, and it resolves every account into one record your team and your agents can act on.

Every account on one record you can trust

Connect your CRM and your inbox, and the account record builds itself from the last 12 months.