The LeanScale Podcast · Episode 119

The Agent Was Never the Hard Part

Kushal Sharma on the data layer underneath AI — what a semantic layer actually is, what a context graph actually does, when a vector database is worth buying, and why almost none of it is an LLM

Kushal Sharma · Head of AI · Circle Hosted by Anthony Enrico
Published Updated 01:07:31 62 min read 12484 words
Executive Summary

The one-paragraph brief, extended

Why this conversation matters — and who should spend the hour.

The episode opens on a vendor call. Census, which Circle used for reverse ETL — pushing enriched warehouse data back into the CRM and the peripheral systems go-to-market teams actually use — told Kushal Sharma that an $18,000 contract was now a $28,000 contract. That is normally where the ops leader begrudgingly signs. Kushal didn't have to, and the reason had nothing to do with negotiation. His team already ran its own orchestration, so in-house scripts fired only after the dbt models they depended on. Reverse ETL, he says, is a basic REST API call you could write in your sleep. Jake Toepel's read: the escape hatch had been built without anyone realising it, and the leverage was instantly gone.

Kushal leads the internal AI organization at Circle, a function he got to build because he had already spent two years merging Circle's data and revenue operations teams. Before that, eight and a half years at Hootsuite, where he led the data organization. Because this one goes deeper than usual, LeanScale CTO Jake Toepel joins Anthony Enrico to take the architecture apart layer by layer.

The semantic layer comes first, and Kushal's definition is deliberately unimpressive: a bunch of text files — a SQL model definition plus a YAML file carrying everything a human needs to understand what is in a table and why it exists. Circle builds it in dbt so the meaning ships with the model in one push. The value is in encoding decisions, not definitions: whether recent churn is weighted more punishingly than churn twelve months ago is a business judgement no NRR column conveys. His litmus test: what a human understands well is likely what an AI will understand, so if you read your semantic data and get no clarity from it, neither will your AI. Underneath sits medallion architecture, bronze raw, silver cleaned, gold joined, with lineage answering what breaks when something upstream changes.

Then the context graph. An LLM is not a store of knowledge; it is a language generator, and the knowledge has to be supplied with the question. RAG did that by hard-coding context; a context graph generates it on the fly, traversing nodes — deals, product usage, transcripts, tickets — to assemble only what the question needs instead of shipping 3,000 pages of knowledge base every time. Almost none of this is an LLM: query handler, graph traversal, packaging, and only then a model. Traditional software: cheap, repeatable, reliable. Jake's addition is that this makes the system model-agnostic, which matters when the gravy train of max plans ends. The sequencing is a maturity ladder: hard-code context in skills while proving value, add a vector database when 20 transcripts a month become 200, consider a graph database when semantics are held together by brute force.

Then what the foundation unlocks. A Slack chatbot built in ten minutes because Notion, Slack and the data were already connected. An AI SDR that has generated close to seven figures, running on a third-party orchestrator with homegrown data and intelligence, working people who fall through the cracks before they become leads. Both hosts caution that messaging still has to be validated by humans. The closing argument is about RevOps itself: capture a stage without its history and your reports die the moment someone changes the value, something Kushal has seen $200–300M companies miss. Get it right and ops heads become product owners with analysts who are now AI engineers. Get it wrong and, as Anthony frames it, you lose relevance and become a support center to sales rather than a go-to-market leading center.

Key Takeaways

15 things worth stealing

The load-bearing ideas, each with the business implication and who should care.

01

Owning your orchestration removed the vendor's leverage before the call ever came

Circle used Census for reverse ETL and it worked well, including on cost, until the renewal conversation moved an $18,000 contract to $28,000. By then the team had already onboarded its own orchestration so in-house data-science scripts and custom workflows ran at the right time and in the right order — a script would not fire before the dbt model it depended on had completed.

Why it matters: Reverse ETL is a basic REST API call. With that foundation, the answer to the price hike was to shift the two or three most important jobs to Airbyte for speed and then move them into their own orchestration. Kushal frames it as deliberate policy: keep the core business pieces under your control so you never suffer vendor lock-in.

RevOps LeadersRevenue ExecutivesFounders
02

A semantic layer is just text — and it ships with the model

Kushal deflates the jargon: a semantic layer is a SQL model definition plus a YAML file containing literally any information a human would need to understand what is in a table, why it exists and where it comes from. Circle builds it in dbt, which lets you model the data while writing the definitions, so one push deploys the data and everything needed to make sense of it.

Why it matters: Without it, that knowledge lives in manually written documents that go obsolete, or tribally in people who leave and take it with them. Kushal's line: if they leave, there goes all your knowledge with them.

RevOps LeadersRevenue Executives
03

Encode the business decisions, not just the column definition

Net retention looks like a no-brainer until you ask whether churn is monthly or annual, and whether the latest churn should be weighted more punishingly than churn from twelve months ago. Those choices have nothing to do with the math and everything to do with how the metric gets applied to your business, and they cannot be teased out of a number sitting in an NRR column.

Why it matters: Saying 'this column contains NRR' is the bare minimum — enough for AI to know where to look. Richer semantics also cover enumerated fields: recording that a stage field holds one of four possible values means never running a query across 300 million records just to find out what is in the data.

RevOps LeadersRevenue Executives
04

The litmus test: if a human can't make sense of it, neither will your AI

Kushal's test for whether semantic data is good enough is deliberately low-tech. What a human understands really well is likely what an AI will understand, so read your own semantic data and see whether you get clarity from it.

Why it matters: This turns an abstract quality bar into something any operator can apply today, without tooling, and without knowing how models work. It is also why the jargon barrier matters — Jake's point is that three-letter acronyms scare teams away from work that is fundamentally readable.

RevOps LeadersFoundersRevenue Executives
05

AI moved documentation from 'always skipped' to default — and moved the bottleneck to review

In the old world, keeping models, docs, dashboards and downstream artifacts up to date meant an analyst could never finish a stakeholder task in the time the stakeholder needed it. Kushal deliberately did not impose that rigor on his team, because a go-to-market data team lives and dies by timeliness. Today a scaffolded skill looks up every related field, document and derivative when a model is produced and flags them on the pull request.

Why it matters: Doing the work is no longer the bottleneck; reviewing it is. Circle now uses AI for a first pass on PR reviews so obvious issues get flagged automatically and human reviewers spend their capacity on the contextual, semantic problems.

RevOps LeadersRevenue Executives
06

Medallion architecture, and letting survivability decide what reaches gold

Bronze is a one-to-one raw copy of source data in the warehouse. Silver pulls the fields you actually need out of JSON blobs, fixes unreadable date formats and cleans column names so a human can start querying. Gold joins tables together into a final artifact — a deal report is useless without the account's vertical, tenure and other context.

Why it matters: Kushal's rule for promotion is survivability: work in silver, produce artifacts, and after thirty reports the five that survive real changes in the business get encoded into gold, because gold demands lineage and history behind it. Lineage — knowing how many derivative tables break when something upstream changes — is what he calls out as the part most teams are missing, and what Jake identifies as why the same column ends up with fifteen definitions.

RevOps LeadersRevenue Executives
07

The CRM is not a data warehouse, and the warehouse is arriving earlier

Jake's observation from LeanScale's smaller clients is that teams use the CRM as a catch-all, dump raw bronze-level data into it, and try to build associations there to produce insights. The result is a mess that takes longer to clean up than it would have taken to structure properly.

Why it matters: The need for a warehouse is moving earlier in the funding journey — what used to feel optional at Series A or B now isn't, because AI tooling is so reliant on context and structure. Anthony adds that usage-based pricing and PLG motions demand the ingestion of more detailed data on their own.

FoundersRevOps LeadersRevenue Executives
08

An LLM is not a store of knowledge

Kushal says this is the thing people who don't know how LLMs work consistently miss. ChatGPT appears to know everything, but it is a language generator: it reads your question, studies the data you supplied with it, compiles an answer and hands it back. It will never know your company's data, because that is not how it works.

Why it matters: Every serious internal AI application is therefore a data-supply problem, not a model problem. RAG systems solved it by taking your query, searching vectorised knowledge-base articles for related topics, and feeding that list plus the question to an LLM API.

RevOps LeadersFoundersRevenue Executives
09

A context graph generates the context on the fly instead of hard-coding it

Previously, context had to be deliberately assembled and hard-coded for every query, which is a lot of engineering time. A context graph treats that as an engineering problem solved at scale: your query is analysed for commonality with everything you know, then the graph is traversed across nodes — the customer's deals, their product usage, transcripts, Zendesk tickets — to decide what to pull.

Why it matters: The data stays where it lives; the graph is a layer on top that knows how it all connects. The payoff is rich, specific answers without boiling the ocean — otherwise you send a 3,000-page knowledge base covering everything from your return policy to how to open an account with every single query, at prohibitive compute cost.

RevOps LeadersRevenue Executives
10

Almost none of this is an LLM — and that is what makes it model-agnostic

Kushal walks the pipeline: a query handler script, a lookup to decide which entity the question pertains to, graph traversal across the nodes, packaging, and only then a call to the model. You are not invoking an LLM at every layer. It is traditional software, which he likes precisely because it is cheap, repeatable and reliable — computers don't make up their minds about things, they just do what they're told.

Why it matters: Jake's extension: people over-complicate this by fixating on which model will be used, when these systems are foundationally model-agnostic. Building that way is how you control your own destiny when the gravy train of max plans giving $4,000 worth of usage ends and billing moves to the API.

RevOps LeadersFoundersRevenue Executives
11

Graduate into the tech curve: hard-code, then vector, then graph

Kushal warns against the eight-month project — by month six the world has changed and some of what you built in month three is obsolete. Instead, hold a strong long-term vision and let maturity dictate the next piece. If your numeric data is a mess, clean that first rather than worrying about vector databases or context graphs.

Why it matters: Two people can spend two or three weeks interviewing stakeholders and hard-coding the context for a weekly revenue meeting into a skill, and AI will run it beautifully. When 20 call transcripts a month become 200, bring in a vector database. When entities grow and you are holding it together with brute force in your semantics, start thinking about a graph database. Don't let it reach 20,000 transcripts before you consider one — but don't worry about it at 20 either.

RevOps LeadersFoundersRevenue Executives
12

Stop throwing your most expensive model at a simple problem

Jake sees teams abandon pipelines as too expensive per run after defaulting to the highest-compute frontier model for work a light model could handle with relative ease — and getting worse results for it. He also notes that call-recording tools teams already pay for often ship sentiment and keyword signals that can be mapped straight into a vector store, cutting compute before a model is ever called.

Why it matters: Kushal reframes the mental model: cheaper models are not dumber models. They run faster, cost less and give you exactly what you need when the problem falls within their purview. Stop ranking models smarter to dumber and start asking what task this one is fine-tuned to do well — a sword isn't better than a kitchen knife.

RevOps LeadersRevenue ExecutivesFounders
13

The AI SDR worked because the foundation and the messaging already existed

Circle's AI SDR has generated close to seven figures in revenue. It runs on a third-party orchestrator platform, but the tech, the data and the intelligence behind it are homegrown. It works the people who fall through the cracks at the top of the funnel before they become leads — researching, analysing, replying to emails, with the goal of getting them to click a scheduler link and book with an AE.

Why it matters: Kushal refuses to trivialise the rest: product, branding and keeping customers happy all drive that outcome, and if nobody wants the brand, no amount of sophisticated AI will make them buy. The hosts' caution is that messaging has to be validated with humans first — you have to know what resonates with your ICP and personas, because AI will fill the gap and you may not like what it chooses to fill it with.

RevOps LeadersRevenue ExecutivesFounders
14

Stage history is the difference between a data org and a prompt-supply team

Kushal's simplest example of operational rigor: if you capture a stage and not its history, your reports become obsolete the moment someone changes the value. Ask how many leads you generated six or twelve months ago and the answer is gone. He has personally seen organisations making $200–300 million in revenue build the field as transient and leave everyone checking old slide decks to find out what last week looked like.

Why it matters: That alone makes moving to AI a non-starter, because all the team can ever do is supply prompts to a chat interface and get an answer back, with nothing foundational that carries historical context and real decision-making power. RevOps has to think about the footprints its operations leave behind — the data those operations generate — not just what it unlocks transactionally for stakeholders today.

RevOps LeadersRevenue Executives
15

Relevance or support center: hire for the ecosystem, not the settings page

Circle's data and revenue operations have always been one team — analysts, analytics engineers and operations people together. Hiring for marketing ops, Kushal read a thousand profiles, wanted to interview about thirty, found three who could speak the language of deep attribution and what scripts on a page actually capture, and hired the one who could hit the ground running. She has since built a lightweight first-party analytics layer on Circle's internal product API that stitches CRM, analytics and self-generated IDs while staying PII-free.

Why it matters: Each ops head now operates as a product owner with analysts who are AI engineers, so a RevOps problem is scoped all the way down to the database tables it touches and every build makes the platform richer. Anthony's cautionary version: without the skills to take control of the data story, RevOps loses relevance and becomes a support center to sales. Kushal's optimistic version: the transition is harder mentally than physically — one of his least technical ops leaders raised their first pull request weeks ago, and the learning curve is three to four focused months at your own pace.

RevOps LeadersRevenue ExecutivesFounders
Frameworks Discussed

9 named models

Every framework Jimmy names, defined and time-stamped.

The Semantic Layer

06:26

A machine- and human-readable description of what is in your data, deployed alongside the data itself: a SQL model definition plus a YAML file carrying the meaning, provenance, calculation choices and enumerated values for every field.

Circle builds it in dbt, which lets you model data while writing its definitions, so one push deploys the model and everything needed to make sense of it. It replaces documentation that goes obsolete and tribal knowledge that walks out with the person who holds it. Historically this needed a dedicated governance team; a small, nimble team could never afford it, so it fell to the wayside.

The Human Readability Litmus Test

12:38

If you read your own semantic data and don't get clarity from it about what the data actually is, neither will your AI.

Kushal's quality bar for semantic data, chosen because it needs no tooling and no understanding of how models work. It reframes AI-readiness as a writing problem any operator can check today.

Medallion Architecture: Bronze, Silver, Gold

19:21

A way of organising a data warehouse in three layers. Bronze is a raw one-to-one copy of source data. Silver is cleaned and sanitised — fields extracted from JSON strings, readable date formats, human-readable column names, only the columns you need. Gold combines multiple tables into the final artifact used for tracking and reporting.

Kushal uses a CRM deal report as the worked example: underneath it sits a deals table, deal stage definitions, and in HubSpot a forced pipeline ID and name. Bronze copies all of it; silver makes it workable; gold joins the deal to its account's vertical and tenure so the report means something. dbt makes it easy to know which derivative tables need updating when something in bronze changes.

Promotion by Survivability

17:07

Living mostly in the silver layer, producing artifacts freely, and letting an artifact's survival through real business change decide whether it earns a place in gold.

Gold demands lineage and history behind it — a claim that this is here to stay and won't be thrown out when the campaign changes in three months. Kushal's ratio: produce thirty reports, and the five that survive get encoded into gold while everything else goes away. Jake later maps the same idea onto a context graph: dump bronze-level material into a personal graph, and promote only proven, strongly-associated content to the company-brain level that gets distributed org-wide.

The Context Graph

28:27

A layer over your existing stores — knowledge articles, call transcripts, support tickets, the relational database — that knows how everything connects and, on demand, traverses those connections to assemble exactly the context a question needs before it reaches the model.

It replaces hard-coded, manually assembled context. Your query is analysed for commonality with all existing knowledge, then the graph traverses nodes — a customer's deals, their product usage, their transcripts and tickets — and packages the relevant subset. The data never moves. The result is rich, specific answers without sending a 3,000-page knowledge base with every query, saving both compute and cost.

None of This Is an LLM

35:09

The recognition that an AI pipeline is mostly traditional software: a query handler script, an entity lookup, graph traversal, packaging — and only at the end a model call.

Kushal's reason for liking it is that traditional software is cheap, repeatable and reliable, because computers don't make up their minds about things. Jake's extension: because the layers are software rather than model calls, the whole system is foundationally model-agnostic, which is what lets a team keep the same performance when pricing or providers change.

Graduating Into the Tech Curve

37:26

Sequencing infrastructure by maturity rather than building it all at once: hard-code context into skills while proving value, add a vector database when unstructured volume grows, and consider a graph database only when your semantics are held together by brute force.

Kushal warns against eight-month builds, where month three's work is obsolete by month six. Instead: hold the long-term vision, then pick the next piece from immediate needs. The staging is explicit — a testing period with no tech investment, an early-adoption period where you buy something fast and expensive per record but cheap monthly, and only then scale. Do the foundational work before the problem gets large, but don't buy for a problem you don't have yet.

Task Fit, Not Smarter and Dumber

50:12

Choosing models by what task they are fine-tuned to do well rather than ranking them on a single intelligence axis.

Cheaper models are not dumber models — they run faster, cost less and give you exactly what you need when the problem falls within their purview, so not using them is a disservice. Kushal's image is that a sword isn't better than a kitchen knife; they are different tools for different purposes. Jake's version of the failure: teams default to the most expensive, highest-compute model for a simplistic problem, then conclude the pipeline costs too much per run.

The Ops Head as Product Owner

1:03:17

Structuring a combined data and revenue operations team so each ops head operates as a product owner, with analysts who are AI engineers and other technical people at their disposal, plus the business strategy and context from the sales or CS leader they partner with.

The whole team rallies behind a RevOps problem and captures it all the way down to the database tables it will affect. The result is that anything new delivers value to the stakeholder and to the platform at the same time, making the platform richer rather than just unlocking a transaction.

Best Quotes

20 lines worth clipping

Pulled verbatim. Copy or share any of them.

“If you take one thing away from this conversation, it's that the agent was never the hard part.”
Anthony Enrico 00:00
“One of the biggest things that I do as a person leading a team with a lot of technical things and a lot of go-to-market data infrastructure is that we keep a lot of our core business pieces under our control so that we never suffer from a vendor lock-in.”
Kushal Sharma 04:11
“You kind of built the escape hatch without even realizing it. The leverage was instantly gone.”
Jake Toepel 05:04
“It's just people that have been there around long enough to know what it is, and if they leave, that there goes all your knowledge with them.”
Kushal Sharma 07:32
“What a human understands really well is likely what an AI will understand. So if you read that semantic data and if you don't get clarity from it as to what it actually is, neither will your AI.”
Kushal Sharma 12:42
“So over time, you don't have to invest in AI tech. You just keep updating your semantic data. You keep updating your skill files and your agents just get smarter and smarter with time.”
Kushal Sharma 24:54
“You can't get to scale and then reverse engineer this whole thing.”
Kushal Sharma 25:08
“A large language model is not a store of knowledge. It appears to be like that because ChatGPT seems to know everything, but it is a language generator.”
Kushal Sharma 26:47
“A lot of this is actually traditional software. You're not invoking an LLM at every layer. All of this is traditional software, which is great news because it is cheap, it is repeatable and it is reliable.”
Kushal Sharma 35:09
“A lot of people over complicate this with what model it's going to be used with when really these systems are foundationally model agnostic. That allows you to control your own destiny when the gravy train of these max plans getting $4,000 worth of usage, that's obviously not gonna be forever.”
Jake Toepel 36:53
“Don't let it go to 20,000 transcripts before you think of a vector database. But you don't wanna worry about it at 20 either.”
Kushal Sharma 39:55
“Think about like a VLOOKUP on steroids, but with text data rather than number data.”
Kushal Sharma 44:14
“People are also throwing the most expensive, highest compute, most complex models at a pretty simplistic problem. This is something that a haiku can handle with relative ease.”
Jake Toepel 49:12
“Instead of us thinking about models in terms of smarter and dumber, we should look at it in terms of what is the kind of task that this is fine-tuned to do really well. It's like a sword isn't better than a kitchen knife. They're just different tools used for different purposes.”
Kushal Sharma 50:12
“My head of data platform and I, like I think three days ago, we built a Slack chat bot in 10 minutes.”
Kushal Sharma 51:21
“I think we have an AI SDR that's generated close to seven figures in revenue.”
Kushal Sharma 53:16
“If you wanna just turn on the Slop Cannon, that's always an option.”
Anthony Enrico 54:47
“If you're capturing a stage and you're not capturing its history, your reports become obsolete as soon as the person has changed the stage value.”
Kushal Sharma 59:35
“Each ops head is now kind of like a product owner. So they have analysts at their disposal who are now AI engineers.”
Kushal Sharma 1:03:17
“If you don't get the skills to take control of the data story, then you're gonna lose a lot of the relevance within the organization and become a support center to sales, not really a go to market leading center.”
Anthony Enrico 1:04:25
Practical Advice

What should you actually do?

The playbook, split by the seat you sit in.

RevOps Leaders

  • Keep the core pieces — orchestration, transformation, the paths your data travels — under your own control, so a vendor price hike is a decision rather than an ultimatum.
  • Build the semantic layer as text that deploys with the model: a SQL definition plus a YAML file describing what the field is, where it comes from and why it is calculated that way.
  • Encode the judgement calls, not just the definitions — how churn is weighted, what a field's enumerated values are — so neither a human nor an agent has to query 300 million rows to find out.
  • Apply the litmus test before you ship: read your own semantic data and see whether you get clarity from it.
  • Organise the warehouse in bronze, silver and gold, and promote to gold only what survives real change in the business.
  • Scaffold a skill that flags every related field, document and derivative on the pull request, then use AI for the first pass on review so humans spend capacity on the semantic issues.
  • Capture stage history, not just current stage — without it, every historical report dies the moment a value changes.
  • Match the model to the task. Check whether the call-recording tool you already pay for gives you sentiment and keyword signals you can load into a vector store instead of paying a frontier model to re-read transcripts.

Revenue Executives

  • Treat the data layer as the thing that decides whether AI works at all — the agent was never the hard part.
  • Fund the foundation while the data is small and the cash is there, instead of paying a team of seven or eight people to clean it up later.
  • Ask what footprints your operations leave behind, not just what they unlock transactionally this quarter.
  • Merge data and revenue operations into one function with analysts, analytics engineers and ops people on the same team.
  • Hire ops leaders who understand the whole ecosystem rather than people who tinker behind a CRM settings page.
  • Validate messaging with humans before pointing an AI SDR at the market — know what resonates with your ICP and personas first.

Founders

  • Stop using the CRM as the catch-all for raw data; the associations you build there become a mess that costs more to unwind than a warehouse would have cost to stand up.
  • Set the foundation early — once everything is scaffolded, adding a source is a ten-minute connector rather than a project.
  • Don't embark on an eight-month build; hold the long-term vision and pick the next piece from immediate needs.
  • While you are still proving value, hard-code context into skills instead of buying tech — two people and two or three weeks of stakeholder interviews can make a weekly revenue review run itself.
  • Buy a vector database when your call volume goes from tens to hundreds a month, not at twenty and not at twenty thousand.
  • Start cheap on infrastructure: a lightweight Postgres instance or a subscription vector store is enough to learn whether the pattern works for you.
AI Takeaways

How AI actually changes GTM

LeanScale's signature read on the AI-in-GTM question this episode wrestles with.

The thesis

The agent was never the hard part. Every capability in this episode — the ten-minute Slack bot, the AI SDR near seven figures, weekly revenue analysis that runs itself — resolves to whether the right context can be assembled and handed to a model, and almost none of that assembly is a model. A semantic layer is text files. A context graph is a traversal over stores that stay where they are. The pipeline is a query handler, an entity lookup, a graph walk, a package, and only then an LLM call. That makes the work cheap, repeatable, reliable and model-agnostic, which is the real hedge when subscription economics change. The operators who could ship AI the week it arrived were the ones who had spent two years making their data mean something.

Agent & automation ideas

  • A documentation skill that fires when a dbt model is produced, looks up every related field, document and derivative, and flags them for approval on the pull request.
  • An AI first-pass PR reviewer that clears obvious issues so human reviewers spend their capacity on contextual and semantic problems.
  • A Slack bot that reads designated channels and answers repeat project-status questions from the Notion docs already connected to the stack.
  • A reusable standardised chat-messenger skill that takes whatever another system outputs, formats it and posts it — so every new bot inherits the communication layer.
  • A weekly revenue agent with stakeholder context hard-coded into a skill, reporting on top-line metrics, new business, churn, product usage, expansions and conversion rates.
  • An AI SDR working the people who fall through the cracks before they become leads — researching, analysing, replying to emails and driving toward a scheduler link and an AE meeting.
  • A sentiment query agent that hits a vector store over MCP to pull only the negative transcripts, then asks whether you want them analysed.
  • A product-direction agent over a nexus of support tickets, bug requests, sales-call challenges and lost-deal reasons, queried in real time with 'what should we build next?'
  • A session-end skill that updates a local context graph so the organisation's knowledge compounds without anyone thinking about it.
Operations Takeaways

By function

The same conversation, filtered for RevOps, pipeline/marketing ops, and customer ops.

Revenue Operations

  • .
  • .
  • .
  • .
  • .
  • .
  • .
  • .
  • .
Metrics Mentioned

The numbers, with context

$18,000 → $28,000
The vendor price hike

The reverse-ETL contract renewal that opens the episode, and the moment Circle's owned orchestration turned a forced signature into a choice.

300 million
Records a semantic layer saves you querying

Kushal's example of why enumerated field values belong in the semantic layer rather than being rediscovered by a distinct-values query every time.

3,000 pages
Knowledge base sent with every query without a context graph

Everything from the return policy to how to open an account, shipped to the LLM each time, at prohibitive compute cost.

$4,000 worth of usage
Max-plan usage allowance

Jake's illustration of the current subscription economics, and why architecture should be model-agnostic before billing moves to the API.

20 → 200 → 20,000 transcripts
Vector database thresholds

Feed 20 straight to the LLM; at 200 bring in a vector database; don't wait until 20,000 to think about one.

10% (200 of 2,000 records)
Relevance rate in an unstructured search

Reading 2,000 transcripts to find the 200 with negative sentiment — what a vector database removes by indexing topics up front.

~$10 a month
Cost of a starter context graph

Jake's point that a lightweight Postgres instance gets a team surprisingly far before any dedicated graph or vector spend.

2–3 weeks for a couple of people
Time to hard-code context for a weekly revenue review

Interviewing stakeholders and coding top-line metrics, new business, churn, product usage, expansions and conversion rates into a skill — a massive efficiency gain with no new tech.

10 minutes
Time to build the Slack chatbot

Built by Kushal and his head of data platform three days before the recording, because Notion, Slack and the data were already connected and a standardised messenger skill already existed.

Close to seven figures
AI SDR revenue

Generated by Circle's AI SDR, which runs on a third-party orchestrator but with homegrown data and intelligence, working people who fall through the cracks at the top of the funnel.

1,000 profiles → 30 interviews → 3 → 1 hire
Marketing ops hiring funnel

Kushal's search for someone who could speak the language of deep attribution; only one could hit the ground running on day one, and the other two would have needed months of training.

$200–300 million
Revenue scale at which stage history is still missed

Organisations Kushal has personally seen build the stage field as transient, leaving no way to answer what the pipeline looked like last week from the data.

3–4 focused months
Learning curve into the new RevOps skill set

Kushal's estimate for moving from Salesforce administration to Markov chains, health-score modelling and data infrastructure, at your own pace and alongside the needs of the business.

~10 minutes
Time to add a new data source once scaffolding exists

Adding a connector to the ETL tool, which is why Kushal argues for building the foundation while the data volume is still small.

Entities

Companies, people & tools mentioned

Auto-extracted and linked into the knowledge graph.

Companies

People

Tools & software

dbtAnalytics Engineering

Circle's semantic layer. Lets the team model data while writing its definitions, so a single push deploys the SQL model plus the YAML file carrying its semantics. Also referenced for lineage — knowing which derivative tables need updating when something in the bronze layer changes — and as the dependency an orchestrated script must wait on before it fires.

CensusReverse ETL

The reverse-ETL tool Circle used to push enriched, combined warehouse data into the CRM and peripheral go-to-market systems. Inexpensive, easy to set up and easy to extend, until the conversation where an $18,000 contract became $28,000.

DagsterData Orchestration

Circle's orchestration layer, transcribed as 'Dexter'/'Daxter'. Already onboarded when the price hike came, running in-house data-science scripts and custom workflows in the right order — holding a script until the dbt model it depends on has completed. Jake calls it the escape hatch built before anyone knew they needed it.

AirbyteData Pipeline

Transcribed as 'AirBite'. Where Circle moved its two or three most important reverse-ETL jobs first, for speed, before shifting them into its own orchestration for granular control.

Amazon RedshiftData Warehouse

Used as the stand-in example of a SQL table when Kushal defines the semantic layer, and again as the relational store sitting alongside a vector database and a graph database in an enterprise architecture.

SupabaseDatabase

Transcribed as 'Superbase'. Jake's example of how cheap a context graph can start — a lightweight Postgres database that can get you really far for about $10 a month.

ObsidianKnowledge Management

Jake's example of building a localized, personal context graph with a lot of success, updated by a skill that runs at the end of every session so it compounds over time.

NotionDocumentation

Where Circle's internal documentation and project status updates live. Already connected to the team's stack, which is why the Slack chatbot that answers project-status questions took ten minutes to build. Also named as the generic example of a knowledge-article store an enterprise graph layer would connect.

SlackTeam Messaging

Where the ten-minute chatbot runs, reading designated channels, replying in others and answering repeat project-status questions from Notion — built on a standardised chat messenger skill reused across all of Circle's bots.

ZendeskCustomer Support / Ticketing

Named as an example of a support-ticket node a context graph would traverse when a question concerns a specific customer.

SalesforceCRM

Named alongside HubSpot as the CRM whose raw tables land one-to-one in a bronze layer, and as the narrow definition of operations Kushal argues people have to grow beyond — from Salesforce administration to health-score modelling and data infrastructure.

HubSpotCRM

Used as the worked example of why raw CRM data needs a medallion layer: a deal report is assembled from several tables, and HubSpot also forces a pipeline ID and name onto the deal.

ChatGPTAI Assistant

The reference point for Kushal's central correction — it seems to know everything, which is why people assume an LLM is a store of knowledge, when it will never know your company's data because that is not how it works.

ClaudeAI Assistant

Named alongside ChatGPT (transcribed as 'cloth'/'cloud') as where teams paste transcripts or supply prompts when they have no foundational data layer underneath. Jake's model-selection point references Haiku as the lighter-weight model that handles simple classification with relative ease.

Model Context Protocol (MCP)AI Integration Protocol

How Kushal says you query a vector store today — you don't go into the database, you connect it to your AI agent over MCP and ask for the records you need.

Google AnalyticsWeb Analytics

The benchmark Kushal used when hiring for marketing operations — candidates who understood how it works and what scripts on a page actually capture. The hire has since built a lightweight version of it on Circle's internal product API, capturing first-party data with CRM, analytics and self-generated IDs, kept PII-free and joined to the product ID on conversion.

Methodologies referenced Reverse ETL under your own controlData lineageRetrieval-augmented generationVectorizing unstructured dataDocumentation as an automated skillFirst-party identity capture
Frequently Asked Questions

Straight answers

Generated from the conversation, marked up for search and AI extraction.

What is a semantic layer in a data stack?

A semantic layer is the description of what your data actually means, stored and deployed alongside the data itself. Kushal Sharma, who leads the internal AI organization at Circle, defines it in the plainest possible terms: it is a bunch of text files — a SQL script that is the model definition, plus a YAML file carrying the text information about that data. It holds literally anything a human would need in order to understand what is in a table and why it exists: where a value comes from, how a calculated column is calculated, which business decisions are encoded in that calculation, and what enumerated values a field can contain. Circle builds it in dbt, so modelling the data and writing its semantics happen together and deploy in a single push. Without it, that knowledge lives in manually written documents that go obsolete, or tribally in people who take it with them when they leave.

What is a context graph and how is it different from RAG?

Retrieval-augmented generation solved a real problem — an LLM has to be handed the knowledge it answers from — but it did so by assembling and hard-coding context for each query, which is expensive engineering time. A context graph generates that context on the fly. Your question is analysed for commonality with all the knowledge in your systems, and the graph traverses nodes — a customer's deals, their product usage, their call transcripts, their support tickets — to decide what to pull and package for the model. The data never moves; the graph is a layer on top that knows how everything connects. The benefit is rich, specific answers without boiling the ocean: instead of shipping a 3,000-page knowledge base with every query, you send only what the question needs, which cuts both compute and cost.

Is an LLM a store of knowledge?

No. Kushal Sharma calls this the thing people who don't know how LLMs work consistently get wrong. A large language model is a language generator, not a knowledge store. ChatGPT appears to know everything, but it is generating an answer to the question you asked by referencing documentation supplied to it. It will never know your company's data, because that is not how it works. So if you want it to answer 'what is my return policy', you have to send your knowledge-based data along with the question; the model reads the prompt, studies the data, compiles the answer and returns it. This is why internal AI systems are fundamentally data-supply problems, and why semantic layers, context graphs and vector databases exist.

When should a company buy a vector database?

When the volume of unstructured data makes brute force wasteful, and not before. Kushal Sharma's guidance is a maturity ladder. If you only have 20 call transcripts a month, feed them all to the LLM — that is fine. When 20 becomes 200, bring in a vector database, because an LLM asked to find negative sentiment across 2,000 records will read all of them to find the 10% that matter, and even free compute burns GPU cycles and time. His framing is that this is like red-lining a car to go to the grocery store. A vector database indexes and catalogs that conversation data by topic up front — think a fuzzy VLOOKUP on text rather than numbers — so only the relevant records ever reach the model. His two bounds: don't let it reach 20,000 transcripts before you think about one, and don't worry about it at 20.

What is medallion architecture?

It is a way of organising a data warehouse in three layers. The bronze layer is a raw, one-to-one copy of source data — if you sync deals from Salesforce or HubSpot, every underlying table lands as it is, so all the raw data is in one place. The silver layer makes it workable: pulling the five fields you need out of a 30-field JSON string into real columns, converting unreadable date formats, cleaning column names, and dropping the fields you don't need. The gold layer joins things together into the final artifact used for tracking and reporting — a deal report is useless without the account's vertical and tenure alongside it. Kushal Sharma adds a discipline on top: gold demands lineage and history behind it, so let artifacts prove themselves in silver first and promote only the ones that survive real business change.

Do AI agents require an LLM at every step?

No, and Kushal Sharma argues that is the best news in the whole architecture. Walking through a context-graph pipeline, he points out that the query handler is a script, the lookup that decides which entity a question pertains to is a query, the graph traversal is traversal, and the packaging is packaging. You are not invoking an LLM at every layer. Only at the end does the model take action on a query that has already been paired with exactly the right knowledge. Traditional software is cheap, repeatable and reliable — computers don't make up their minds, they do what they're told. It also means these systems are foundationally model-agnostic, which is what lets a team keep performance steady when pricing changes or a better model arrives.

Why do cheaper AI models often beat expensive ones in production?

Because most production tasks are narrower than the model thrown at them. Teams default to the most expensive, highest-compute, most complex model for a simplistic problem, then conclude the pipeline costs too much per run — when a lighter model would have handled it with relative ease and sometimes produced better results. Kushal Sharma's reframe is that cheaper models are not dumber models: they run faster, cost less, and give you exactly what you need when the problem falls within their purview, so refusing to use them is a disservice. Stop ranking models on a single smarter-to-dumber axis and ask instead what task a given model is fine-tuned to do well. A sword isn't better than a kitchen knife; they are different tools for different purposes.

What should RevOps teams do to stay relevant as AI arrives?

Take control of the data story. Anthony Enrico's cautionary framing on the episode is that a RevOps team without those skills loses relevance inside the organization and becomes a support center to sales rather than a go-to-market leading center. Kushal Sharma's structural answer is to integrate RevOps with data as step one — one team of analysts, analytics engineers and operations people — and then run each ops head as a product owner with technical people at their disposal, so a problem is scoped all the way down to the database tables it touches. It starts with rigor as basic as capturing stage history, not just current stage. His optimistic footnote: the transition is harder mentally than physically. One of his least technical operations leaders raised their first pull request a few weeks before recording, and he estimates three to four months of focused, self-paced learning to make the move.

Full Transcript

The whole conversation

Broken into chapters, searchable, verbatim from the audio. Speakers inferred (not diarized).

00:00Cold open and intro

0:00 If you take one thing away from this conversation,

0:03 it's that the agent was never the hard part.

0:06 Joining me today is Kishal Sharma,

0:09 who leads the internal AI organization at Circle,

0:12 a function he got to build because he had already spent

0:15 two years unifying Circle's data and

0:17 revenue operation teams under a single mandate.

0:21 Before Circle, Kishal spent roughly eight and a half years at

0:24 Hootsuite where he ended up leading

0:26 the company's data organization,

0:28 we're also throwing the most expensive, highest compute,

0:33 most complex models at a pretty simplistic problem.

0:37 Cheaper models actually,

0:38 people think of them as dumber models.

0:41 I framed it as an optimistic point,

0:44 maybe I'll make it a little bit more cautionary.

0:46 I think if you do want to still add

0:48 a positive point to that cautionary piece,

0:51 it is that that transition may sound

0:54 harder on paper than it actually is.

0:56 The really interesting thing about that is, to your point,

0:58 it is just traditional software,

1:00 and I think a lot of people over-complicate this

1:02 with what model it's going to be used with,

1:05 when really these systems are foundationally model-agnostic.

1:08 That allows you to control your own destiny to your point when

1:11 the gravy train of these max plans,

1:14 that's obviously not going to be forever.

1:15 [MUSIC]

01:30The vendor call: an $18K contract becomes $28K

1:30 >> Tell me about the day Census told you

1:33 the $18,000 contract was now a $28,000 contract,

1:38 because your reaction was not the reaction most ops leaders have.

1:43 >> Sure. Thank you. We were using it for basically reverse ETL stuff.

1:49 We were using it for making sure that any data that we were processing in

1:54 the warehouse and enriching it,

1:57 combining it with other data sources,

1:59 we can make that available to the CRM and

2:01 other peripheral systems that the go-to-market teams actually use.

2:05 That was really what we were using it for.

2:07 It was relatively inexpensive,

2:09 easy to set up, easy to extend.

2:12 All of the typical things that work well with a tool like that,

2:15 were working well for us,

2:16 including the cost until that conversation where we

2:19 realized that the cost is going to jump quite a bit significantly.

2:23 At that point, basically,

2:27 what we had also done was set up a really strong rigor

2:32 in terms of how we orchestrated our stuff in the background.

2:35 We had already onboarded Dexter by the time to make sure

2:39 that any data science scripts that we were developing in-house

2:43 or any custom workflows we were developing in-house,

2:47 we could deploy them and we could run them at the right time.

2:50 If it depends on a specific DBT model to be

2:53 run before the script can have the data it needs,

2:56 you want to make sure that it doesn't

2:57 fire before the DBT model is completed.

3:00 We already had that ecosystem in place.

3:03 At that point, when we were put into

3:05 that position where a lot of vendors unfortunately do today,

3:08 where they have you because they have all your data,

3:11 they have all your systems,

3:12 and they know you can't disentangle yourself from their ecosystem,

3:16 and then they will show you a price hike that you have

3:18 to basically just begrudgingly accept.

3:20 Fortunately, we had laid the foundations

3:22 to never be in that position with a vendor,

3:25 and so when that conversation came along,

3:27 we basically said, "You know what?

3:28 We have two options. We have AirBite,

3:30 which also does diverse ETL,

3:31 and we have Dexter where we can write our own scripts."

3:34 Diverse ETL is actually really straightforward.

3:36 It's a basic REST API call that especially with AI,

3:42 like you could write it in your sleep,

3:44 and so basically we just went that route.

3:46 We took our two or three most important things,

3:49 first shifted them to AirBite to make it fast,

3:52 and then after that,

3:53 we basically moved them over to Dexter,

3:56 and that way we were able to control them more granularly

4:00 and moved on from there,

4:01 so that was basically our response.

4:03 One of the biggest things that I do as a person leading a team

4:07 with a lot of technical things

4:09 and a lot of go-to-market data infrastructure

4:11 is that we keep a lot of our core business pieces

4:15 under our control

4:16 so that we never suffer from a vendor lock-in basically.

4:21 I think it's really smart.

04:23Owning the data layer means owning your leverage

4:23 Jake, from your perspective,

4:24 have a lot of things changed in the tech environment

4:27 that enable it to be easier,

4:30 more efficient to own the data layer,

4:34 own the data model so you don't get locked in

4:36 to those hostage situations with vendors?

4:39 Yeah, totally.

4:41 It's a completely different landscape now, I would say,

4:43 and Kishall nailed it.

4:44 You can write these Lambda functions,

4:47 these reverse ETL pipelines with one line.

4:50 You just say, "I want to move data from here to here,"

4:53 and the AI takes it from there.

4:55 It's pretty amazing how when you control the data,

4:58 you really control your own destiny,

5:00 and the thing that I love about this is that,

5:02 to your point, Daxter went in first,

5:04 and you kind of built the escape hatch

5:06 without even realizing it.

5:07 The leverage was instantly gone,

5:09 and I think that's what we're seeing

5:10 is more and more companies are wanting to control their data,

5:14 and the data is really where these products thrive.

5:16 They need it to be successful,

5:18 but ultimately, it's so table stakes now

5:22 to get your data in the right format, right shape,

5:25 and where you want it,

5:26 that what you layer on top of that

5:27 and what you even can build on top of that

5:29 is really endless nowadays.

5:31 So it's a really interesting market,

5:33 and I think SaaS vendors are really trying to figure out

5:36 what their mode is now

5:37 when customers are so protective of their data

5:40 and aren't just going to leverage themselves

5:42 to a specific system.

5:43 They need to have that flexibility,

5:45 and the data needs to be available and usable

5:48 in all of these downstream systems.

5:51 >>Yeah, and I think something that's interesting

5:53 is this data layer, the data model has become,

5:55 it's always been important,

5:57 but now, it is crucial if you want to start layering in

6:01 agents, agentic workflows, any AI type of automations.

06:06What a semantic layer actually is

6:06 Kaushal, some terms that get thrown around,

6:08 I'm curious how you define them,

6:10 semantic layer, context graph.

6:13 For people who might not be as technical,

6:16 how would you explain it

6:18 and then get into the core of what those two terms mean?

6:22 >>Totally, yeah, thank you.

6:24 So let's start with semantic layer

6:26 because that's something that has existed for a while,

6:29 and essentially, what that means is,

6:32 you have a table of data, let's say,

6:36 let's say, Redshift, for example,

6:37 or you could do SQL table of any kind.

6:40 Now, you have fields in it that each field,

6:42 basically, each of which means something, right?

6:45 You have an MRR column,

6:47 you have a column for something else,

6:50 like all of these values,

6:51 and oftentimes, the description

6:54 and the intro of that column

6:58 to a person who doesn't know about it

7:00 is a lot more involved than just saying,

7:02 "Oh, this is where we captured this value,"

7:05 because sometimes a column may contain a value

7:07 that is actually calculated,

7:09 and different companies may have a different way

7:12 of calculating that.

7:13 So all of this really rich information

7:16 about what this data really is,

7:18 most of the time in companies,

7:20 it lives either in documents that are manually written,

7:23 which means that they're also going

7:25 to become obsolete pretty quickly,

7:27 and managing them can become a nightmare,

7:31 or it lives tribally.

7:32 It's just people that have been there around

7:34 long enough to know what it is,

7:36 and if they leave,

7:37 that there goes all your knowledge with them, right?

7:40 So that kind of fragmentation

7:41 is really what the semantic layer

7:43 is kind of designed to solve for,

7:46 and we use DBT for that,

7:48 and DBT sort of does it in a really nice way

7:51 in that DBT allows you to model your data

7:56 while you're building the definitions

7:58 and the semantics for it.

8:00 And so in one push,

8:02 when you send out that model to be deployed,

8:06 you're sending out all the data necessary

8:08 to make sense of it with it, right?

8:10 Now, in the past,

8:11 this, at a company that was really serious about it,

8:15 required a specific mandate to do this.

8:18 It could be somebody's job to do this.

8:20 There would be like a governance team

8:22 that was deployed whose only job was to make sure

8:24 that this information was up to date, right?

8:26 And in a small nimble team,

8:28 you can't really afford that all the time.

8:30 So this is the stuff that always fell to the wayside,

8:33 but with the advent of AI,

8:35 that's the first thing we tackle on the inside for ourselves,

8:38 is the tool we built for ourselves, right?

8:40 For the tool makers,

8:41 which is that now when we build a DBT model,

8:44 we have an actual skill associated with the AI workflow

8:48 that works with it,

8:49 so that the documentation happens along with it.

8:51 So all the administrative tasks, you don't have to do it.

8:54 And this semantic layer essentially contains

8:56 literally any information that a human being would need

9:00 to understand what is in the table

9:02 and why it exists, where it comes from,

9:04 anything you can put into it

9:05 that is a value for that data, it contains that.

9:08 And what does it look like in reality?

9:10 It's a bunch of text files, basically,

9:12 whether that's a .yaml format or something else,

9:16 it's essentially that.

9:17 So there is a SQL script that is your model definition,

9:19 and then there is a .yaml file with it

9:21 that has text information or all of the basic semantic data

9:25 about this information that you wanna deploy.

9:27 So that basically is a semantic model.

9:30 Now-- - Let me--

9:32 - Yeah, go ahead.

9:33 - Well, I wanna hang on there for just a little bit.

9:34 So it makes a ton of sense.

9:36 So somebody who might not be as familiar with it,

9:39 it's literally just text to context for an agent to read

9:44 and then go take a look at the data

9:45 and then make sense of it based on those text files.

9:48 What are some good go-to-market type of use cases

9:54 where that would be super important?

9:56 - Oh, absolutely.

9:57 Like, I mean, things like even, for example,

10:01 net retention rate, let's take that as an example

10:03 as a metric, that metric is on the face of it,

10:09 it is a no-brainer that it's basically the dollars

10:12 that you earned from a cohort of customers a year ago,

10:15 how many of those dollars have you retained today, right?

10:18 Contraction reduces the amount of the dollars,

10:20 expansion increases that, and so on and so forth.

10:23 And churn takes away from it, that's that.

10:27 However, depending on how your business is structured,

10:30 you could have churn monthly, you could have churn annually,

10:33 you could have churn anywhere in between.

10:35 And if you're trying to aggregate a monthly churn model

10:39 into an annualized number,

10:41 then do you weight the latest churn

10:43 that has happened more highly than your previous churn?

10:46 See, because how you decide that

10:47 depends on what it does for your business, right?

10:51 If in your business, the latest churn

10:53 has a bigger indicator of what's happening today

10:56 versus what used to happen 12 months ago,

10:58 and if that's why you want that to be more valuable,

11:00 you would make that churn more punishing to your model

11:03 than you would the earlier stuff, right?

11:06 So you're encoding all of these really important decisions

11:09 that have nothing to do with the actual calculation

11:11 and the math of it.

11:12 It has more to do with how you're gonna apply it

11:13 to your business, right?

11:15 And so that nuance cannot always be teased apart

11:19 from just a number that is in an NRR column.

11:22 So you can encode a lot of data

11:25 that is relevant to you in that semantic layer

11:29 that isn't just this column contains the NRR value.

11:32 That's the bare minimum you could do with it.

11:34 So at least AI or whoever else is looking at it

11:36 know where to look for it, but you can make it richer

11:38 to understand where it actually comes from.

11:41 And then there could be fields

11:42 that have enumerated values, for example.

11:45 Let's say that you have a field

11:46 that contains one of four possible values

11:48 whether that's a pipeline stage or a lead stage

11:51 or a customer's potential place in their journey

11:54 like stage in their journey.

11:56 And let's say if you have 300 million records

11:59 for a query to run that entire column

12:03 to find out how many distinct values it has

12:06 that's not always a thing you want your system

12:08 to be doing every single time.

12:10 So if your semantic layer can already encode that

12:13 and say this field contains one of four possible values

12:16 and you keep it up to date,

12:17 you never have to run a query to find out

12:18 what's actually in the data, right?

12:21 So it's stuff like this that becomes super useful

12:26 when you're talking about thousands of tables

12:29 in your database, millions and millions and millions

12:31 of rows of records and so on and so forth.

12:35The litmus test — if a human can't read it, neither can your AI

12:35 And the beauty of it today especially is that

12:38 the litmus test of whether this is good

12:40 is actually really cool

12:42 because what a human understands really well

12:44 is likely what an AI will understand.

12:46 So if you read that semantic data

12:49 and if you don't get clarity from it

12:51 as to what it actually is, neither will your AI.

12:53 So it's a great litmus test to set it up

12:55 in a way that AI agents actually do

12:57 what they're supposed to do

12:58 which is can a human looking at this make sense of it?

13:02 - That's a really interesting way to break it down.

13:04 I actually really like that.

13:05 A very practical use case and easy to understand.

13:08 I think what we see a lot of the times is

13:10 some of these concepts almost become a barrier to entry

13:13 because of the jargon and some of the three letter acronyms

13:16 that are thrown around

13:17 and people just get a little overwhelmed

13:18 and intimidated by that.

13:20 But when you say, hey, this is really just for the AI

13:23 to understand and for you to understand exactly

13:26 what's going on behind the scenes,

13:28 it's a really easy pill I think for most teams to swallow

13:30 and that's why I love the approach that you all have taken

13:33 of unifying those two teams.

13:35 I always use the Uber example which is a bad one to use

13:38 but they've boiled it down and said,

13:40 we're a logistics company.

13:42 Regardless of if it's a person, food, package,

13:44 it doesn't matter, they're logistics.

13:46 They're not a ride sharing company or a taxi company.

13:49 And I think most RevOps orgs are kind of missing

13:52 that portion to distill it down and say we are a data org.

13:56 That's what drives the revenue insights.

13:58 That's what drives RevOps, GTM ops.

14:01 So I think you guys are approaching it

14:02 in a really unique way and breaking it down

14:05 like that just makes it much more consumable

14:07 for someone who's not in the weeds of these DVTs and rags

14:10 and all the MCPs and three letter acronyms galore.

14:14 - Yeah.

14:15 Well, it's definitely a barrier for me sometimes.

14:17 I'm like half technical on TV sometimes.

14:22Protecting the context layer as the organization scales

14:22 So I think one question I have is

14:24 as an organization is scaling,

14:27 is there anything that you need to be cognizant of

14:30 to protect and maintain this layer of context?

14:33 Are there certain tools that you should be using

14:36 that are different?

14:37 Are there certain processes you should be following

14:38 that are different as your organization grows

14:41 and gets more complex over time?

14:43 - Yeah, and I think that's what I would say

14:46 that the biggest impact of AI has been in this area.

14:49 Because if you look at the entire workflow

14:52 of developing a metric, let's say tracking for a metric,

14:57 in the old world, what you would have is a person

15:01 writing the query, saving the model,

15:05 making sure that it works.

15:06 But while you're iterating on it,

15:08 your model has to be changed multiple times

15:10 because you may use one thing that you don't need later

15:12 or you may not use something that you need later, whatever.

15:16 And then developing the actual output

15:19 that the business is gonna use,

15:20 whether that's a dashboard or something.

15:22 In that entire space, one person can brute force all of this.

15:27 They can produce all the 10s of different documents

15:29 that are gonna be needed for all of that.

15:31 But then they have to move on to the next project.

15:35 And if anything changes, and if they have to,

15:37 like over let's say three or four years of their history,

15:40 three or four years of their history,

15:42 if they have to keep making sure that everything they created

15:45 in the last three years stays up to date,

15:47 they will never be able to finish a stakeholder task

15:50 in the time that the stakeholder needs the answer, right?

15:53 They'll be like, oh, I gotta do these 20 things

15:56 before I can give you yours,

15:57 because this will change eight other artifacts

15:59 that I've got that I need to now go and update

16:00 all the documentation for.

16:02 I need to inform all of those people.

16:04 I need to update those dashboards and metrics.

16:06 And so most of the time what happens is you just don't do it.

16:09 You just do the bare minimum, you move on,

16:11 and you hope that one day when you grow,

16:13 you'll have enough money to hire a team

16:14 to clean up the mess, right?

16:18 That was the world we lived in.

16:20 And so in that world,

16:21 actually my approach used to be slightly different.

16:24 I did not put these types of onerous practices on my team

16:28 because I knew that as a team

16:30 that services go to market priorities,

16:32 we live and die by the timeliness of the value

16:35 we provide to our stakeholders, right?

16:37 And so we would do the bare minimum that we could

16:39 for it to make sense to us and to onboard somebody else,

16:43 but we wouldn't go all the way.

16:45 And in fact, in the medallion architecture,

16:47 our goal there would be the least developed one,

16:50 because to get something to gold,

16:52 it has to have lineage behind it.

16:53 It has to have history behind it

16:55 for it to say that this is here to stay.

16:57 This isn't gonna change in three months

16:58 when your campaign changes,

17:00 because otherwise all the work you did

17:02 architecting that entirety of the landscape

17:05 is gonna be thrown out, right?

17:07 So you mostly lived in the silver layers

17:09 and you kept producing artifacts

17:11 and you allowed survivability of that artifact

17:14 to define what goes into the gold layer.

17:16 So after you produce 30 different reports,

17:18 if five survive the changes in the business that happens,

17:21 they can get encoded into the gold layer

17:23 'cause they're here to stay, but everything else goes away.

17:25 So coming back to your question, Anthony,

17:28 today, however, what we can do

17:30 is we can scaffold skills in the right way

17:33 to make this just a matter of automation, essentially.

17:36 So when you produce a model,

17:38 the skill looks up all the related fields, documents,

17:41 and derivatives of it that need to be looked at.

17:44 As you raise your PR with your, let's say,

17:47 the person who's managing the merges into the Git repo,

17:51 it will flag all of those things.

17:53 And it will tell you, should I update this?

17:54 Should I update that? Should I update that?

17:55 And then you go, you review it and you say,

17:57 "Yep, go ahead and update it."

17:58 So today, actually doing the work

18:00 is not a bottleneck anymore.

18:02 All of that administrative work has been,

18:04 if you are intelligent about how you set up your workflow,

18:08 you can have all of that automated.

18:10 The bottleneck is on the review side.

18:12 The person reviewing it now gets a lot of code

18:15 generated by AI every single day.

18:18 They're the ones where they're dealing with capacity issues.

18:21 So now they're also using AI to create a first pass

18:24 of PR reviews that can happen more automatically.

18:29 And obvious issues can be highlighted.

18:31 And then the more contextual, more semantic issues,

18:36 that is the one that the human reviewers then gets to see.

18:38 So today, I think it's been the biggest gift of AI

18:43 that documentation and sort of,

18:47 I suppose a broad term, project management of these things,

18:51 that can be automated.

18:53 If you can spend a little bit of time

18:54 configuring your tooling the right way,

18:57 you can basically, it gets done by default, essentially.

19:01 It's not a thing you have to dedicate time to do separately.

19:04Medallion architecture: bronze, silver, gold

19:04 - Makes a ton of sense.

19:06 I wanna take a quick step back to make sure

19:08 we clarify some of the terms you threw out.

19:11 When you were talking about the medallion system

19:14 of data layers, can you walk me through

19:16 the difference between gold, silver, bronze,

19:18 and what that might mean for a team?

19:20 - Absolutely, yeah.

19:21 It's just a way of organizing your database, basically,

19:23 and your data warehouse.

19:25 So bronze layer is typically at the bottom of it

19:28 where essentially it's just a one-to-one raw mapping

19:32 of your data.

19:33 So if you're syncing data from a CRM, Salesforce, or HubSpot,

19:37 their reporting aggregates a lot of data.

19:39 But if you see underneath, let's say even a deal data,

19:42 like a deal sort of a report,

19:44 is a combination of five different tables, for example.

19:47 And I'm just throwing the number five out just like that,

19:50 but basically you have a basic deals ID,

19:52 then you have deal stage definitions and their IDs,

19:55 then you'll have, if it's HubSpot,

19:56 they also force you to attach a pipeline to it,

19:58 basically, so there's a pipeline ID and name,

20:01 and then it aggregates all of these

20:03 to show you what the actual information is

20:04 that you're asking for.

20:06 A bronze layer basically takes a copy of all of that

20:09 and puts it into your warehouse as it is.

20:11 That's what it is.

20:12 It's only meant for you to have all the raw data

20:14 in one place, you don't have to go anywhere else.

20:17 The silver layer is where you take individual bronze tables

20:21 and you go, this field is a JSON string

20:24 of 30 different fields, out of which I need five.

20:28 So I'm gonna pull those five out

20:29 and make them specific columns so I can access them faster.

20:33 Or this date time format is not human readable.

20:36 It's written in a certain way.

20:37 I wanna change it and put it in a certain other way,

20:39 or what have you.

20:41 So what you do is you take that bronze raw data layer

20:43 and you make it a little more readable.

20:46 You clean up the column names

20:48 so that they are also more human readable,

20:50 and you kinda like sanitize it.

20:53 And then you also say, well, the raw layer

20:56 has 40 different fields, do I need all of that?

20:59 I maybe only need 15.

21:00 So you'll take 15 and you bring those into it.

21:02 So the silver layer becomes the clean place

21:05 where a human can go and start working

21:06 with a baseline query to say,

21:08 what do I wanna actually look in this table

21:10 and how do I connect it to the others?

21:12 The gold layer is the one where now

21:15 you're bringing everything together.

21:16 So our deal report is kinda useless on its own

21:20 if you don't know what account it's from, right?

21:22 But on the deal table, you only have account IDs.

21:25 You need other data on that account to bring it in.

21:27 Maybe they're vertical.

21:28 Maybe when you last, they became your customer,

21:32 if they're an existing customer, all of that stuff.

21:35 So in your gold layer, what you do

21:36 is you combine multiple tables to produce a final artifact

21:41 that gives you everything you're gonna need

21:43 for your tracking, reporting, and what have you, right?

21:46 So that's how it's a good way of organizing it

21:49 and DBT or tools like that make it really easy

21:52 for you to do it, where if something

21:56 in the bronze layer changes, how many derivative tables

21:59 need to be updated to make sure?

22:00 That is what is called lineage

22:02 in the data analytics engineering world, right?

22:05 So the lineage is basically that if something changes here,

22:08 how many other things are tied to it

22:10 that need to be updated to make sure things don't break?

22:12 And that is also then captured.

22:13 So this medallion architecture allows you

22:15 to really organize everything.

22:18Lineage, and why the CRM is not your warehouse

22:18 - Yeah, it's a great call out on the lineage.

22:20 I think that's the big missing piece for a lot of folks.

22:22 And what we see is then you have 15 definitions

22:25 of what this column means across many different schemas

22:27 or tables.

22:29 The other, I think, interesting thing here

22:31 that we see from a lot of our smaller clients

22:33 is using the CRM as kind of your catch all

22:37 and just dumping all of that bronze raw data in

22:39 and trying to make those associations in the CRM

22:43 to produce the insights that they're looking for.

22:46 What we find is it creates a huge mess

22:48 that actually takes longer to clean up.

22:50 So it seems like the need for a data warehouse

22:52 is actually moving a little bit earlier

22:55 in the company's life cycle or funding journey

22:57 where previously, you know, series A, maybe series B even,

23:00 it was like, I don't know if we need that yet

23:02 or if our team needs access to that.

23:05 That's just for product data for right now.

23:07 But I think what we're seeing is especially to your point

23:09 with the rise of AI and some of these tools

23:12 that are so reliant on the context and the information

23:15 that structuring that early

23:16 and kind of building this foundation from the very beginning

23:19 will save you a ton of time down the road

23:23 versus going back and trying to make sense of everything

23:26 from the CRM and this kind of hodgepodge mess

23:28 that you end up building.

23:30 - Well, and a lot of companies too,

23:33 they're moving to a monetization model

23:35 that's tied to usage-based pricing.

23:38 They're moving towards PLG motions

23:40 that requires the ingestion of more detailed data.

23:44 So even just the business model of a lot of these SaaS

23:47 or tech companies is moving towards go-to-market motions.

23:52 That really demands a better data structure.

23:56 - Absolutely, yeah.

23:58 And you know, if a company has any hopes

24:03 of scaling their output and their insights

24:07 in a way that they won't need to throw more bodies

24:09 at the problem, then getting ahead of this problem,

24:13 making sure your data infrastructure is clean

24:15 and loaded up when you don't have a ton of data

24:18 'cause then you can brute force it as well if you need to.

24:21 Doing it ahead of time when you have the cash to spare

24:24 and you don't have that much data

24:25 that would take like a team of seven, eight people

24:27 to clean things up, that's actually the right time to do it.

24:30 'Cause if you scaffold everything in its place,

24:33 adding more things is just a matter

24:35 of adding a new connector to your ETL tool, right?

24:38 That's like a 10 minute thing, right?

24:41 But having this in place allows you to set the foundation

24:44 where you've set up your AI to reference

24:47 all of these things.

24:48 It already is trained to look at semantic data.

24:51 It already understands your metrics that are evolving.

24:54 So over time, you don't have to invest in AI tech.

24:58 You just keep updating your semantic data.

25:00 You keep updating your skill files

25:02 and your agents just get smarter and smarter with time.

25:05 So it just, like that's how you set yourself up for scale.

25:08 You can't get to scale

25:10 and then reverse engineer this whole thing.

25:12 That becomes a far more painful transition

25:15 than if you were to like get ahead of it a little bit

25:17 when the problem isn't as big

25:18 and set the foundations right.

25:21 - Yeah, definitely cause much more to do or renovation

25:25 than to just build it right from the beginning.

25:27Context graphs from first principles

25:27 I wanna get into the other major area

25:31 that helps you get even deeper leverage out of your AI.

25:35 That's the context graph.

25:37 Similarly, how we walk through the semantic layer.

25:39 Can you walk through kind of first principles,

25:41 what context graph is and then the tools you use to manage

25:46 and what it can actually unlock for your agents?

25:50 - Yeah, totally.

25:51 I think a context graph is one of those sort of emerging

25:55 pieces of tech where, you know, basically,

25:58 if you go back to the early days of chat GPT,

26:02 where any third-party system that is, let's say,

26:06 a customer service agent or any kind of a chat bot

26:10 that isn't a generic chat bot that comes out of the box,

26:13 you know, like chat GPT,

26:16 the thing that people didn't understand in the beginning,

26:18 is I fielded these questions, which is like, you know,

26:20 why do you need a separate chat bot

26:22 if you can just ask chat GPT?

26:23 I was like, well, you need it

26:24 because chat GPT doesn't know your company's data

26:26 and it will never know your company's data

26:28 because that's not how it works, right?

26:30 What you have to do is you use it to generate answers,

26:34 but what it generates answers on

26:37 needs to be supplied with the question, right?

26:41 So this is the part that I think a lot of people

26:43 don't fully understand,

26:44 especially those who don't quite know how LLMs work.

26:47 Is that a large language model is not a store of knowledge?

26:50 It appears to be like that

26:53 because chat GPT seems to know everything,

26:55 but it is a language generator, right?

26:59 It is generating an answer for you

27:02 for the question you've asked,

27:04 but it generates that by referencing other documentation

27:07 that has the answers you're looking for.

27:10 So if you wanna know what is my return policy,

27:14 along with that question,

27:15 you send it all of your knowledge-based data

27:18 that contains your policies outlined.

27:21 Then it will read your question,

27:24 understand what that question means,

27:26 look up the knowledge base to understand what the answer is,

27:29 compile that answer for you and give it back to you, right?

27:33 That's usually how LLMs work.

27:34 You send the prompt, you send the data.

27:37 It reads the prompt, it studies the data,

27:39 generates the answer and gives you the answer back.

27:42 And for that type of a pipeline,

27:45 you had these systems before called reg,

27:48 like resource augmented generative.

27:50 And these were these types of systems

27:52 where the form in which you were typing your chat query

27:55 was not actually the chat box input.

27:57 It was an input into a software system

28:00 where it would take your query,

28:02 it would then go through its list of knowledge-based articles

28:05 that would typically be stored in like a vectorized format

28:08 in like a vector database or something.

28:10 And it would pull topics that were related to that question

28:12 just as a loose search.

28:14 Then that entire list of topics with the question

28:16 would be fed to an LLM API, like open APIs,

28:20 open AI's API, right?

28:22 And based off of that, you would get the response back

28:24 and that response then get sent back to the user.

28:27 So now coming back to what is a context graph?

28:32 So a context graph does this thing

28:35 that I talked about in the backend,

28:36 which is taking your question

28:39 and then finding all of your knowledge that exists

28:42 that could have anything to do with that question

28:43 and compiling it as context to be given

28:47 with your query to the large language model,

28:49 that context generation in the past

28:52 had to be deliberately done manually,

28:56 where you had to set up the context and hard code that

28:59 and supply it.

29:00 So that's a lot of engineering time to do that.

29:02 And what a context graph as a system allows you to do

29:05 is generate that on the fly.

29:07 So now people are realizing

29:08 that this is an engineering problem

29:11 that could be solved at scale instead of forcing every user

29:15 to generate their entire set of contexts

29:17 for every query and hard coding it.

29:19 What you can do is you can make that more nimble.

29:22 And so what a context graph would typically do

29:25 is without getting into the technical details

29:28 of how these systems work,

29:29 basically your query is just analyzed

29:34 for its commonality with all of the knowledge

29:37 that exists in your system.

29:39 And it will traverse that graph across the different nodes.

29:41 So if you're asking about a customer,

29:43 one of the nodes could be all of the deals

29:44 they've done with you.

29:46 Another node could be all of their product usage.

29:48 Another node could be some other transcripts

29:50 or Zendesk tickets or customer service inquiries

29:53 or what have you that you have on that customer.

29:56 These are all of the different nodes

29:57 of data about that customer.

29:59 The context graph as a tool can traverse these nodes

30:02 to find out what it should pull

30:04 that would be a good knowledge base

30:07 for this question you've asked.

30:09 And then it would supply that to the LLM on the fly basically.

30:12 So that's in essence what context graphs do.

30:15 What it allows the AI to basically do

30:17 is give you very rich answers

30:19 that are very specific to your use case

30:22 without boiling the ocean every single time.

30:25 So if you didn't do this,

30:27 you could have a knowledge graph base of 3000 pages

30:31 that contains topics all the way from

30:34 your return policy to information for the teams on something

30:39 or how to open an account to a million different things

30:43 that have nothing to do with your query.

30:45 You would be sending all of that every single time

30:48 for the LLM to make sense of your question.

30:50 The compute power that you would spend on to do that

30:53 would be just prohibitively expensive for a lot of people,

30:56 especially now that pricing is no longer cheap, right?

31:00 Models are getting expensive

31:01 and companies now gotta make money.

31:03 So we are all hooked on it

31:05 because we were given all of that for free.

31:08 Now it's gonna start costing a lot of money.

31:09 So things like context graph will help save the effort

31:13 both in terms of the compute power you need to do it

31:16 and also in terms of the cost you're paying

31:18 to answer that question.

31:20 All of that gets optimized

31:21 if you use something like context graph properly.

31:25Starting cheap: Obsidian, Postgres and a company brain

31:25 - Yeah, it's so interesting and I think you nailed it.

31:28 I liken this almost a little bit to the medallion model

31:32 where you can set this up early

31:34 and it's not so much of a barrier to entry anymore

31:37 to start kind of just organizing some of this data.

31:40 Things like Obsidian I've seen a ton of success with

31:42 in terms of just building like a localized context graph

31:45 but even just setting it up in something like Superbase

31:47 which it's a lightweight Postgres database

31:50 but it can get you really, really far for $10 a month.

31:54 So I think the barrier to entry for setting some of this up

31:58 is much lower.

31:59 I think where it gets a little bit interesting

32:01 is when you start to think about it

32:03 through the medallion architecture lens

32:05 where I say, hey, my personal context graph,

32:07 I'm fine just dumping bronze level data into that

32:11 but we wanna have also layers on this

32:14 where only the things that are proven out

32:16 have the strong associations

32:18 are making it kind of to that company brain level

32:21 where things are then distributed around an entire org.

32:24 So I really liked that comparison

32:25 and also just the lack of resources

32:28 that you truly need to start standing this up.

32:31 To your point, it can be just in a skill

32:32 that runs at the end of every session,

32:34 updates everything, you don't even have to think about it

32:37 and then that just compounds over time.

32:40 - Exactly.

32:41Context graphs at enterprise scale: relational, vector and graph

32:41 - What does the architecture look like

32:43 at an enterprise level?

32:44 When you have enterprise level data, complex organization,

32:49 what are the tools that you use,

32:50 the processes that you put in place?

32:52 How do you actually manage something

32:54 like a context graph at scale?

32:56 - Totally.

32:57 I mean, I think that is a question

32:59 basically for companies that are,

33:01 this is one of those things

33:03 that you actually don't have to get ahead of

33:05 because it's truly a thing about efficiency

33:09 rather than about efficacy, right?

33:11 You can have efficacy by hard coding skills

33:14 like Jake was talking about.

33:15 You can have efficacy

33:16 by just making your semantic layer really rich

33:19 to make sure that the specific queries you're asking

33:23 are kind of answered well.

33:24 So for example, even as we talked about

33:28 redshift, which is like a typical

33:31 relational database type of store,

33:33 then we talked about a vector database.

33:35 We didn't go into the details of what that is,

33:37 but basically that's where you're essentially

33:39 making sense and indexing all of your unstructured data,

33:43 your call transcripts and what have you.

33:45 The third thing would be, let's say a graph database.

33:49 That's another type of database

33:50 that basically stores entities and their relationships, right?

33:55 So you could imagine having, let's say one storage layer

34:00 for all of your knowledge articles.

34:03 Like if people use Notion or whatever

34:04 for internal documentation,

34:06 let's say that you have one place

34:08 where all of that text is stored.

34:10 You have another place, let's say,

34:11 where all of your call transcripts are stored.

34:13 You have another place where, let's say,

34:15 you have all of your support inquiries and tickets stored.

34:19 And then you have your RDBMS

34:20 where you have all your data stored.

34:22 What a graph layer then would do is make sure

34:26 that these are all connected.

34:30 So you wouldn't put all of this data in one place.

34:32 It still lives where it lives,

34:34 but you have this sort of a layer on top of it

34:37 that knows how all of it connects, right?

34:40 And then on demand, when a question comes in,

34:43 it traverses those connections to say,

34:45 for this thing, I need this data, I need these docs,

34:49 I need those tickets, I need these transcripts.

34:51 Now we can formulate an actual answer

34:54 based off of this data.

34:55 So it would essentially be this layer

34:57 that would traverse that,

34:58 and it would be a pipeline, basically.

35:00 So it wouldn't be like one system that does this.

35:02 You have your query handler, right?

35:05 Let's say that could be a script.

35:08 And this is the cool part about this.

35:09"None of this is an LLM" — the software underneath

35:09 A lot of this is actually traditional software.

35:12 None of this is like,

35:13 you're not invoking an LLM at every layer, right?

35:16 All of this is traditional software, which is great news

35:19 because it is cheap, it is repeatable and it is reliable.

35:22 Those are the three biggest things about basic software,

35:25 which is like computers don't make up their minds

35:28 about things, they just do what they're told, right?

35:30 And so you have your question that comes in immediately.

35:34 Now, let's say a query runs on top of that question

35:37 to say, this question pertains to which entity

35:40 for my graph system to kind of look at.

35:43 Okay, this is the entry.

35:45 Now the graph traversal will now go through

35:47 all of your different nodes that exist and say,

35:51 these are all of the things that are important.

35:53 Pull that, package it, send it back to another script.

35:56 Now that's the packages your query

35:57 and the knowledge pushes it to the LLM.

36:00 Now the LLM is taking action, right?

36:03 And then you get the answer.

36:05 So it's like you're architecting your entire pipeline

36:09 in a way that allows you to control how the work is done

36:13 before the LLM even gets into it, basically.

36:16 And I just gave like one very rough example of this,

36:18 but I'm sure that people more experienced smarter than me

36:21 are doing it even in more cooler ways.

36:24 But yeah, that's kind of like the essence of it.

36:26 Like you just architect that entire system

36:28 where you have the start to the end of the process

36:31 and you know where a script will get invoked,

36:33 where matches will be made and data will be pulled from.

36:36 And then finally, what will be sent

36:37 to your large language model to get the answer out of it.

36:41 - Yeah, you nailed it.

36:42 And I think the really interesting thing about that is,

36:45 to your point, it is just traditional software.

36:47 And I think a lot of people over complicate this

36:50 with what model it's going to be used with

36:52 when really these systems

36:53 are foundationally model agnostic.

36:56 That allows you to control your own destiny to your point

36:58 when the gravy train of these max plans

37:01 getting $4,000 worth of usage,

37:04 that's obviously not gonna be forever.

37:06 So building these in a way that they're very agnostic

37:08 to what you layer on top of them

37:10 and still get the same performance,

37:12 I think to your point, it's a huge efficiency play.

37:15 And it also allows these teams to control their own destiny

37:18 and not really have that wolf in the hen house experience

37:21 of, oh my gosh, when they switched to API billing,

37:23 I'm not gonna be able to afford any of this.

37:26The maturity ladder: when to actually buy a vector database

37:26 - Yeah, totally.

37:28 And you know, the way that you would even go

37:30 about developing this is that you never wanna embark

37:32 on an eight month project

37:33 where you're gonna do all of this at once.

37:35 Because by the time month number six arrives,

37:37 like world level have changed.

37:39 Some tech would have gone obsolete,

37:41 some new stuff would have arrived

37:42 that makes some of the stuff you did

37:44 in month number three useless or onerous or whatever.

37:48 Instead, what you wanna do is have this long-term vision

37:51 where you know these are the high level pieces

37:54 you're gonna need, have that vision

37:56 really, really strongly developed.

37:58 But then as you get to the individual steps,

38:00 look at what are your immediate needs and priorities

38:03 and where do they fit in that vision, right?

38:06 So if today even your numeric data is a mess,

38:09 I wouldn't worry about vector databases or context graphs,

38:13 I would clean up the numeric data right away.

38:15 'Cause at a minimum, your number crunching becomes real.

38:18 When you look at your data, you know you can trust it, right?

38:21 And context, sure, you can start providing

38:23 the missing context in skills

38:25 by hard coding it for smaller use cases, right?

38:28 That already is a massive efficiency gain.

38:30 Like for example, if you do a weekly meeting

38:32 that looks at all of your revenue across the board

38:34 where you talk about all your top line metrics,

38:36 your new business acquisition, your churn,

38:38 your product usage, your, you know, like expansions

38:41 and conversion rates and what have you,

38:43 that is not that much of a work.

38:45 Like you could, a couple of people could spend like

38:47 two to three weeks of time writing all of that context down

38:50 by just interviewing stakeholders

38:52 and coding it properly in the skill, connecting it and saying

38:55 this belongs to these fields and this is the data.

38:58 Your AI will do a marvelous job of running this

39:01 on a weekly basis for you

39:02 and giving you some really great insights.

39:04 So you start with that.

39:05 But then when the skill piece of it

39:07 starts to get more complex, when you're like,

39:09 oh, I wanna now start adding unstructured data to it.

39:12 I wanna now start analyzing these things.

39:15 At that time, let's say if you still only have 20 calls

39:18 a month, if you're a small business,

39:20 those 20 calls don't need to be vectorized.

39:22 You can feed all of that to the LLM, that's fine.

39:24 But then when that 20 becomes 200,

39:26 okay, let's bring in a vector database now, right?

39:29 And now all of a sudden you've added one more layer to it.

39:31 And then as your entities grow and it becomes more complex

39:34 and now you're realizing that you're holding all of this

39:36 together with a lot of brute force in your semantics,

39:39 at that point, maybe start thinking about a graph database.

39:42 So it's like, you can graduate, but it doesn't,

39:45 so it may sound a little bit like I'm contradicting myself

39:48 where I talked about doing a lot of foundational work,

39:50 but the point here is that you do the foundational work

39:53 before the problem gets really large.

39:55 So don't let it go to 20,000 transcripts

39:58 before you think of a vector database, right?

40:00 But you don't wanna worry about it at 20 either.

40:02 You, when you're still in the middle of understanding

40:06 how useful this is for you, at that point,

40:08 don't invest in tech, just do the work, right?

40:11 Hard-coded in the skill, understand the value it drives.

40:14 Once the value is established, then start doing maybe,

40:18 maybe get like a cloud-based store

40:19 where you pay more per record, that's okay.

40:22 You know, VVA is a great example for a vector database.

40:24 Like if you wanna just start with it,

40:26 not invest too much in it to understand

40:27 whether it's gonna work for you,

40:29 just get a subscription for that.

40:30 Get your vectorized data into it, start using it,

40:34 see how well it goes.

40:36 Now you know you wanna scale it,

40:38 then invest in something that may be on like,

40:40 in your own cloud or what have you, right?

40:42 So it's almost like you're letting your maturity dictate

40:47 what piece of that grand vision you're gonna put next,

40:50 and you're doing it in a way that you start doing it

40:53 before it becomes a problem.

40:54 You have a testing period,

40:56 then you have a early product adoption period

40:59 where you buy something fast but expensive

41:02 on a per record basis but it's still cheap

41:04 on a monthly aggregate level,

41:06 and then you talk about scale.

41:08 That way you've graduated into the curve

41:10 of that tech as you need it, right?

41:12 - That's the way we think about it here too.

41:14 It's making stage fit decisions.

41:17 You don't need to play business at a certain point

41:19 in your company's maturity.

41:21 But like you said, it may be being half a step ahead.

41:24 So hey, you know the scale's coming,

41:26 start working on these things.

41:28 I'd like to dive deeper into vector database

41:31 and just talking about why it's important,

41:35 when you mentioned, hey, at 20 transcripts

41:38 it's not a big deal at 2000,

41:40 it's gonna be really important.

41:41 So talking through those factors

41:44 that start to make it important and why,

41:48 and then technically how to go about implementing

41:52 or anything that teams should be thinking about

41:54 before deploying.

41:56 - Yeah, absolutely.

41:57 It's such a great topic actually, right?

41:59 Because when you have unstructured data

42:04 that contains a goldmine

42:06 of so much sort of qualitative information

42:09 that you may not be getting in your numbers,

42:12 it's almost like a missed opportunity

42:14 to not be able to leverage that.

42:16 LLMs and AI agents make it really easy to do that today.

42:20 But in the early days,

42:22 most of the time what people will do

42:24 is that they will say, here are my transcripts,

42:27 go chat GPT or cloth or whatever,

42:30 find me all of the calls I've had

42:35 where a negative sentiment was expressed.

42:38 So typically what an LLM would do

42:40 is it would just read everything you gave it, right?

42:43 But unbeknownst to you in the backend,

42:46 it's actually sampling

42:47 because it's not gonna go through 2000 records, right?

42:50 So if the record count is small,

42:54 like I said, 20 transcripts,

42:55 it'll read all of it, after reading each one,

42:58 it will understand if it was actually negative

43:01 in any parts of it.

43:02 And if it wasn't negative, it'll discard.

43:04 And if it was negative based on what you asked for,

43:06 it'll keep that.

43:07 Then it will compile the list of transcripts

43:10 that specifically had negative commentary,

43:12 then it will analyze those and then give you the answer.

43:15 Now, imagine having to do that every single time

43:18 and your records aren't 20, but 200, 2000.

43:22 And out of which,

43:23 let's say that the incidence of negative commentary

43:25 ends up 10% of the time.

43:27 Now what you're doing is you're going through 2000 records

43:30 to find only 200 that actually are relevant

43:32 to your inquiry, right?

43:34 That is expensive.

43:36 That kind of compute,

43:37 even if somebody gives it to you for free today,

43:39 the GPU cycles are what they are.

43:42 It takes how long it takes for the system to do this.

43:45 So at a minimum, you're just like,

43:48 you're essentially red lining

43:50 to go to a grocery store, essentially.

43:51 Like if you do the car analogy, right?

43:54 It's not a thing that you want to do

43:57 or is sustainable in the longterm.

43:59 So what a vector database would let you do

44:01 is that it would automatically catalog and index

44:05 all of this conversation data

44:08 into all of the different topics that it could pertain to.

44:12 Which means that when you,

44:14 so think about like a VLOOKUP on steroids,

44:17 but with text data rather than number data.

44:22 So it's like basically if you say negative,

44:26 it will do like a fuzzy VLOOKUP on all 2000 records

44:30 to find the 200 that already have been tokenized

44:33 and sorry, like by been sort of,

44:35 they've been classified in topics

44:37 to say that these had negative.

44:39 And those 200 only are the ones that go to the LLM.

44:43 So the first part of it is not LLM.

44:45 It's very cheap, like traditional software.

44:47 And vector databases have been around forever by the way.

44:50 So, you know, it's just a simple query.

44:54 It runs a query.

44:56Vectorizing transcripts: themes, not keywords

44:56 - Are there certain techniques

44:57 that are important to deploy?

44:59 So imagine you're ingesting, I mean, transcripts are,

45:02 I don't call it gold, I called it diamonds

45:04 because they're so valuable

45:06 and you could produce so much from a transcript.

45:09 And it's even changed the way I even approach

45:12 building out a plan or building out playbooks here.

45:15 It's now turned into like,

45:16 hey, let's just set up an hour meeting,

45:17 talk at each other for a while

45:20 and then we can pretty much build anything we want.

45:21 But as you're ingesting, let's say thousands of transcripts,

45:25 let's use a go-to-market example of sales calls,

45:27 customer calls, are there certain techniques

45:31 that you would be doing ahead of time to say,

45:33 hey, I wanna make sure I'm vectoring these pieces of data

45:37 or this unstructured data by these ways

45:39 as they're coming in?

45:40 Do you have to do some of that upfront?

45:42 What does that actual process look like?

45:44 - Yeah, it's like so loading, you know,

45:46 loading data into your vector store is a little bit different

45:50 than loading data into something like a relational database,

45:52 like a SQL database, 'cause SQL database is just one-to-one.

45:55 This is your table.

45:56 If you look at it in a CSV file or a spreadsheet,

45:58 it looks like a table.

45:59 And in a database, it looks exactly like that table,

46:02 basically, right?

46:03 In a vector database, it's a little bit different

46:05 because for vectorizing, quote unquote, as a process,

46:10 what you're really doing is you are taking this piece of,

46:16 you know, text that you have,

46:19 and what you're doing is

46:20 you're kind of natural language processing it,

46:22 where you're attaching sort of topics to it,

46:26 or you're sort of vectorizing it in a way

46:29 that you can look it up by themes,

46:31 rather than looking it up by exactly the keyword.

46:34 So if you have, you know, you're talking about,

46:37 let's say, cats, and if the word feline is used,

46:42 the system needs to be smart enough

46:43 to know these two are connected, right?

46:47 And then, you know,

46:49 if you're talking about a stuffed toy that was a cat,

46:52 that is a slightly more distant relationship

46:55 than cat and feline right away, right?

46:57 But then in a certain context,

46:58 if you talk about a toy store and a cat, right?

47:02 Then the stuffed toy becomes more relevant.

47:04 And so there are actually, so in the old world

47:09 where vector technology was new,

47:11 you had to define how you wanna do this mapping.

47:14 Today, you have established models

47:17 that will do that for you.

47:18 So it's actually a lot easier.

47:20 And think about it this way too.

47:21 Like if you were gonna use an AI thing,

47:23 AI tool that generates not text for you,

47:26 but let's say query, SQL queries for you,

47:29 now you're vectorizing SQL queries

47:31 for it to be able to generate those.

47:32 And queries, the words don't come next to each other

47:35 in the same way they do a natural language.

47:38 So a sentence written in a natural language

47:40 versus a sentence written in a SQL query,

47:43 the proximity of words is completely different.

47:44 So there's a very different mapping

47:46 that needs to be applied when you're vectorizing code

47:49 versus when you're vectorizing language, right?

47:51 And so these models exist already.

47:53 It sounds more complex and it is complex

47:56 when you try to design something like that,

47:58 but its usage is not complex.

48:00 You can just take a mapping and then you can use that

48:03 and then you can vectorize it.

48:04 And then your vector collection

48:06 will just basically be populated.

48:08 And when it's populated, then you can start querying it.

48:11 And these days you don't even have to go

48:13 into the vector database to query it.

48:14 You just connected via an MCP with your AI agent

48:17 and you just say, this is my vector store,

48:19 give me all the topics that were negative.

48:21 And it will just run a query

48:23 instead of doing the heavy lifting that I told earlier,

48:26 where it read all 2000 transcripts

48:28 to find the 200 that were bad.

48:29 Now what it's gonna do is it'll just run a query

48:31 against the vector store, pull the 200 that you need

48:34 and it will say, I got the records,

48:35 what do you wanna do now?

48:36 Do you want me to analyze it?

48:38 Yes, go ahead.

48:38 And then you're done.

48:41Stop throwing your best model at a simple problem

48:41 Yeah, two things I think are super interesting

48:43 in this realm are one, we see a lot of people in the market

48:48 leaving meat on the bone when it comes

48:50 to the call recording tools that they already pay for.

48:52 I mean, a lot of these already come with certain keys

48:55 on sentiment analysis or certain keywords

48:57 that are called out within the calls,

49:00 which you can easily then map into your vector

49:02 and kind of reduce some of that compute costs

49:04 that we're seeing.

49:05 And then the other thing is we hear a lot of like,

49:07 oh, we tried to set up this pipeline,

49:10 but it just got really expensive.

49:12 And I think people are also throwing the most expensive,

49:16 highest compute, most complex models

49:19 at a pretty simplistic problem, right?

49:21 This is something that a haiku can handle with relative ease.

49:25 And I think that's still something

49:26 that most people just default to let me use the best.

49:29 And then they say, I can't support this,

49:31 this is costing too much per run.

49:33 When in reality, they'd get maybe even better results

49:36 from a lighter weight model.

49:39 - So true, like, I mean,

49:40 this is becoming such an important problem now

49:42 when it comes to cost control

49:44 and like using your resources more efficiently,

49:46 that it absolutely is right.

49:48 'Cause cheaper models actually,

49:50 you know, people think of them as dumber models.

49:53 And yeah, in some point of view and a paradigm that is true,

49:57 but actually in other ways, they're better

50:00 because they run faster, they cost less money,

50:04 and they just give you exactly what you need

50:06 if your problem is something that falls within its purview,

50:09 right?

50:10 So to not use it and actually doing yourself a disservice

50:12 because instead of us thinking about models

50:15 in terms of smarter and dumber,

50:17 we should look at it in terms of like,

50:18 what is the kind of task that this is fine-tuned

50:20 to do really well?

50:22 Make sure you give it that task, right?

50:24 It's not worse or better.

50:26 It's like a sword isn't better than a kitchen knife, right?

50:30 They're just different tools used for different purposes.

50:32 And I think that's the kind of mentality

50:34 we gotta approach this with.

50:36 - Well, you should see how Jake cuts his stake.

50:39 You might have a different opinion on that, but.

50:42 No, I think it's gonna become increasingly more important

50:46 as the frontier models get more expensive,

50:47 like you mentioned.

50:48 So you're gonna have to start to segregate,

50:49 okay, these tasks assigned to these models,

50:52 those tasks assigned to these models,

50:54 let's vectorize the database

50:55 so our models can run efficiently

50:57 when they're trying to pull data in certain scenarios.

51:00What it unlocks: a Slack bot in 10 minutes

51:00 I'm curious, we've talked a lot about the foundation,

51:06 the infrastructure, all the piping.

51:09 Once you have this, what does this unlock for a GTM team?

51:16 - Absolutely incredible value.

51:18 Like I'll give you one example.

51:21 My head of data platform and I, like I think three days ago,

51:25 we built a Slack chat bot in 10 minutes.

51:31 And it went from a question that somebody asked saying,

51:37 hey, can I buy this product?

51:38 'Cause it integrates in Slack, it reads all the messages,

51:41 and it gives replies, and it connects with Notion.

51:44 And when I'm working on my projects,

51:46 I have these things that have a status update on each of them,

51:49 but it's a real pain to answer the same question

51:52 a million times over in Slack and to keep an eye on all of it.

51:54 Can a bot just do it for me?

51:56 And because we already have all of our Notion docs connected,

52:00 we have our Slack connected, we have our data connected,

52:04 the exact kind of scaffolding that we just talked about.

52:07 We just went into our chat bot builder,

52:10 that's like a simple system we have on top of this.

52:13 We said, okay, what is the bot's name here?

52:15 What channels should it read from these channels?

52:17 What channels should we reply to here?

52:19 Or like, how should we reply this way, whatever?

52:22 What data does it need access to, these Notion source?

52:26 What other skill does it need?

52:27 Oh, we have a standardized chat messenger.

52:31 So that way, this messenger doesn't care

52:33 what you're messaging.

52:34 It just looks at the format and generates the response

52:37 and handles the communication.

52:40 So the data that the other system is giving,

52:42 it just takes it and then it turns it into a message,

52:44 post it on Slack.

52:45 So we just reuse that for all the bots, basically.

52:48 Decided all of that, deployed it,

52:51 and now we have a chat bot.

52:53 - That's insane.

52:54 That's insane.

52:55 What about some other, just like the table stakes stuff

52:59 GTM teams have to do, things like forecasting,

53:01 managing customer health, or forecasting churn,

53:05 or moving things through a pipeline.

53:08 How has this infrastructure helped unlock

53:11 the basics that have to get done every day,

53:13 that every company has to do?

53:15 - Absolutely.

53:16The AI SDR that generated close to seven figures

53:16 I think we have an AISDR that's generated

53:22 close to seven figures in revenue.

53:26 - Wow.

53:28 - Right?

53:29 - Homebuilt, homegrown, AISDR.

53:31 - Well, I mean, we used a orchestrator platform,

53:33 third party platform for it,

53:35 but all the tech, all the data, all the stuff,

53:39 the intelligence behind it is homegrown, basically.

53:42 Yeah, these are just the people that fall

53:45 through the cracks in the funnel,

53:46 top line before they become leads.

53:48 And these are the people that the AISDR researches,

53:50 analyzes, communicates with them,

53:54 gets them to, replies to their emails,

53:57 and with the ultimate goal of being to help them,

54:01 click on that scheduler link to book a call with an AE.

54:05 And then from those, we get meetings that come out of it.

54:07 And from those meetings, we basically get sales out of it.

54:10 And yeah, it's a pretty insane number that,

54:14 we're able to see come out the other end based on that.

54:17 - Yeah, that is, yeah.

54:19 - That's amazing.

54:20 And I think to your point,

54:22 what we see a lot of or get a lot of requests is,

54:25 I really want AI to help me kind of get my message out there

54:29 to these prospects, to the market.

54:32 But we have to kind of remind folks,

54:33 especially when they're going to use a third party solution

54:35 and bring on one of these AISDR platforms,

54:38 is you have to validate the messaging with humans.

54:41 You have to know what resonates with your ICP,

54:43 with your different personas.

54:45 The AI is not gonna figure that out for you.

54:47 I mean, if you wanna just turn on the Slop Cannon,

54:49 that's always an option.

54:51 But I love the way that you guys approached it to say,

54:53 hey, we've built out all the foundation,

54:55 we have all the context.

54:56 We know our ICP, we have the messaging already.

54:59 Now we just need to build more channels

55:01 to get it out there and to blast it

55:02 and to be available 24/7.

55:04 And I think that's the correct approach.

55:06 A lot of teams put the cart a little bit before the horse

55:10 and hope that AI will fill that gap.

55:12 And what we've seen is, it will,

55:14 you just might not like the results.

55:16 - Exactly.

55:18 Like what it chooses to fill in the gap

55:19 may not be what you wanted it to fill.

55:21 (laughing)

55:22 Yeah, 100% agree.

55:24 And that's obviously a very dramatic outcome

55:27 that I don't know if everybody can replicate on demand.

55:33 The opportunity has to exist

55:34 for that kind of returns to happen, of course.

55:36 And that depends on a lot more than just your tech setup.

55:40 If it's a brand that nobody wants to buy,

55:42 you can put the most sophisticated AI engine

55:45 in the backend and generate as many leads as you can,

55:47 but it's not gonna buy anybody.

55:47 So it's not gonna make anybody buy it, right?

55:50 So obviously, the investment we do in our product,

55:53 the investment we do in our branding

55:56 and in keeping our customers happy,

55:57 all of that plays a role in the actual revenue outcome.

56:01 So I'm not gonna trivialize all of that

56:02 and put it all on, let's say this SDR agent

56:05 that we built that's doing magic, right?

56:07 But what it does do for you

56:11 is all other things being in place,

56:13 it unlocks the kind of efficiency

56:15 where you really can hit the ground running

56:18 with a lot of things and go from conception to deployment

56:22 in a remarkably short period of time.

56:24 And that's really it.

56:25 So understanding health scores,

56:28 prepping for your customer calls

56:30 in a way that gives real value to your customers,

56:33 focusing on the people that actually need your help

56:35 more than everybody else in your book of business

56:38 and targeting engagement in a very specific way.

56:44 A design team wants to know what to build next.

56:47 If they have a nexus point that contains the synthesis

56:50 of all your support tickets, all your bug requests,

56:54 all your sales call challenges that you faced,

56:57 all of the reasons why somebody shows

56:59 not to be your customer.

57:01 If they have that ability to query that real time

57:04 and say, what should I build next?

57:06 I mean, the kind of power that unlocks for you

57:08 is just phenomenal, right?

57:09 People used to pay insane amounts of money

57:13 to learn this stuff like to be able to do this right.

57:16 And today you can basically all just have it

57:18 if you just get a little bit ahead of the problem.

57:22From RevOps to AI: rebranding the team

57:22 - Kishal, one thing I think is so interesting.

57:26 You rebranding your team from RevOps to AI.

57:32 I'm seeing a huge trend where RevOps is in the driver's seat.

57:38 It's your ticket to take if you want

57:41 to own the AI strategy internally

57:44 because you're working in all the systems already.

57:46 You have so much of the business context

57:48 plus you tend to have a lot of the technical skills

57:51 on the team.

57:53 Walk me through that transition

57:55 and what do you think it means for RevOps teams in SAS today?

58:01 - Yeah, I think none of this happens

58:05 without RevOps becoming tightly integrated with data.

58:10 And by data, I don't just mean that

58:13 they run a lot of reports and they build a lot of spreadsheets

58:17 and they do all the planning,

58:18 which is all good work that is acquired and people do.

58:21 It's absolutely not a knock on that.

58:23 But to truly unlock value,

58:28 it's like, so in my mind,

58:30 this is how the analogy works, right?

58:32 Is that RevOps, a lot of times when it is divorced from data

58:36 gets looked upon as a system that lets people do stuff, right?

58:42 Oh, a salesperson wants to send a contract out, fine.

58:45 You integrate it with some kind of CPQ tool

58:47 or a contract management system

58:48 and then you kind of push out the capability.

58:51 They wanna know certain things.

58:53 All right, you create a new field in your CRM

58:55 so you can capture that data

58:56 and you can start reporting on it.

58:58 A lot of times that becomes the focus.

59:01 Like what am I unlocking today,

59:03 transactionally for my stakeholders?

59:06 But very little thought goes into the footprints

59:09 that operation leaves behind, right?

59:12 And that footprint is the data that your operation generates.

59:16 It can generate good data or it can generate crappy data

59:19 that just clouds up your pipes

59:21 and doesn't tell you anything valuable.

59:23 And that's where, right?

59:24 The rigor with data comes into play.

59:26 I'll give you the simplest of examples

59:27 because it's just so fundamental.

59:29 Stage matching, right?

59:31Stage history, and the footprints your operations leave behind

59:31 Or stage mapping, whether it's a deal, a lead or whatever.

59:35 If you're capturing a stage

59:37 and you're not capturing its history,

59:39 your reports become obsolete

59:41 as soon as the person has changed the stage value, right?

59:44 So if you wanna know how many leads we generated

59:46 six months ago, 12 months ago,

59:48 all of that data has gone

59:49 because the stages have changed, right?

59:52 So if somebody says,

59:53 I wanna know at what stage a lead is at any point in time,

59:57 I kid you not, it may sound really stupid to do this,

1:00:00 but I have personally seen at large organizations

1:00:04 making two to $300 million in revenue

1:00:07 where people did not think about that.

1:00:09 They just built a field that is transient.

1:00:12 And in that field, you get the value that it is today

1:00:15 and it leaves you hanging in terms of like,

1:00:17 what did I have last week?

1:00:18 Nobody knows.

1:00:19 If you made a slide deck about it in your weekly meeting,

1:00:22 you go through your slide decks

1:00:23 to know what it was yesterday or last week,

1:00:25 but you don't know it from your data, right?

1:00:28 So it's stuff like that.

1:00:30 Like that already makes integration

1:00:33 and moving to AI a non-starter

1:00:34 because at that point, all your team will ever do

1:00:37 is be a group of people that just supplies prompts

1:00:40 in chat, GPT or cloud

1:00:42 and gets an answer out of it that way,

1:00:44 but doesn't have anything foundational

1:00:45 that is long lasting

1:00:47 and can give you that historical context

1:00:49 and real decision-making power.

1:00:52 So integrating data, I would say it was a step one.

1:00:55 And that was the genius here, especially at Circle,

1:00:57 which is that we saw that early on

1:01:00 and our data and DevOps team has always been together.

1:01:02 It is one team where we have analysts,

1:01:05 we have analytics engineers

1:01:06 and we have operations people.

1:01:09 That was one.

1:01:10 And then what we started to do was we...

1:01:13 So now this is where I think subjectivity comes into play,

1:01:16 but for us, when I was specifically hiring,

1:01:20Hiring: a thousand profiles, three interviews, one hire

1:01:20 I did not hire an ops person

1:01:23 that only lived in their ops system.

1:01:27 So like, I'll give you an example.

1:01:29 When I was hiring for marketing operations,

1:01:31 I kid you not, I went through a thousand profiles,

1:01:34 like actually a thousand profiles.

1:01:36 Like I read through all of them,

1:01:38 out of which I found maybe 30

1:01:40 that I wanted to actually interview.

1:01:42 And out of the 30, I found three

1:01:45 that actually could speak the language

1:01:47 and understand the systems that we were talking about

1:01:49 in terms of the deep attribution methods,

1:01:55 how that even works.

1:01:56 How does Google Analytics work?

1:01:57 What are the scripts running on your page

1:01:59 that captured the data that you want

1:02:01 for marketing attribution?

1:02:03 And then out of those,

1:02:04 I found one who could hit the ground running on day one.

1:02:07 The other two would have required

1:02:08 at least a few months of training.

1:02:10 And that is the person that we ended up hiring.

1:02:12 And today she's built a lightweight version

1:02:15 of Google Analytics for us

1:02:17 using our product API that is internal.

1:02:20 And based off of that,

1:02:22 we're actually capturing first party data

1:02:24 that at the time of your visit captures all IDs,

1:02:28 including your CRM ID, your Google Analytics ID,

1:02:33 an ID that we generate ourselves

1:02:35 to keep a track record of it and keep it PII free.

1:02:38 And then as soon as the conversion happens

1:02:40 attaching the product ID to it.

1:02:42 Like we have a lightweight segment, if you will,

1:02:45 that we are using to do this.

1:02:47 And that's the kind of power

1:02:50 that we put into our operational rigor.

1:02:52 So when I hired operational leaders,

1:02:55 I hired leaders that aren't just people

1:02:57 that tinker behind a CRM setting page

1:02:59 and just create new fields or workflows.

1:03:01 These are people that fundamentally understand

1:03:04 what is the rest of the ecosystem need,

1:03:06 what is that ecosystem and what does it need to look like

1:03:09 for our operations to be not brittle

1:03:12 and scalable and what have you.

1:03:15 So we brought that together.

1:03:17 And then each ops head is now kind of like a product owner.

1:03:21 So they have analysts at their disposal

1:03:23 who are now AI engineers.

1:03:25 And they have other technical people at their disposal.

1:03:28 And they have the business strategy and the context

1:03:30 with the sales leader they work with

1:03:32 or the CS leader they work with.

1:03:34 And then the whole team kind of rallies behind.

1:03:36 So when they talk about a problem

1:03:38 that we wanna solve in RevOps,

1:03:40 all the way down to the database tables

1:03:42 that will be affected by it,

1:03:43 there is a full capture of that.

1:03:45 And then it gets built that way

1:03:47 so that anything that we do that is new,

1:03:50 it doesn't just deliver value to our stakeholder,

1:03:52 it also delivers value to our platform

1:03:54 and makes the platform richer, basically.

1:03:59 - Well, it's a huge transition.

1:04:02 And I think laying down the foundations

1:04:05 and the fundamentals as you have,

1:04:07 put you in the position to lead the circle

1:04:10 through the AI transformation,

1:04:12 especially on the go to market front.

1:04:13 And I think anybody who's listening to the sets

1:04:15 in the RevOps world needs to know,

1:04:18 I'd actually, I framed it as an optimistic point.

1:04:23 Maybe I'll make it a little bit more cautionary.

1:04:25 If you don't get the skills

1:04:27Relevance or support center: the choice facing RevOps

1:04:27 to take control of the data story,

1:04:30 then you're gonna lose a lot of the relevance

1:04:33 within the organization and become a support center to sales,

1:04:37 not really a go to market leading center.

1:04:40 So I think everything you've done is the roadmap

1:04:44 that ahead of RevOps should be following,

1:04:47 structuring their team in the new world of AI.

1:04:51 - Absolutely.

1:04:52 And I think if you do wanna still add a positive

1:04:54 sort of point to that cautionary piece,

1:04:57 it is that that transition may sound harder on paper

1:05:00 than it actually is, right?

1:05:02 It's today the best part about all of the toolkit available,

1:05:06 the entire toolkit available to us is that

1:05:09 you can ask it to teach you anything you don't know, right?

1:05:13 We don't have to approach AI

1:05:14 in the way we used to approach software,

1:05:16 which is, oh, I don't have experience with it.

1:05:18 I don't know how to use it.

1:05:19 Well, that too, you can ask that AI

1:05:21 and it will tell you how to use it, right?

1:05:23 (laughing)

1:05:24 - Just keep running.

1:05:26 - Exactly.

1:05:27 One of our least technical operations leaders

1:05:31 basically just raised their first pull request

1:05:36 like two or three weeks ago.

1:05:38 And the transition was harder mentally

1:05:41 than it was actually physically to do.

1:05:44 So the point is that, yes, there is a lot to learn,

1:05:47 there is a lot of change and there is a lot to do,

1:05:50 but these are actually great opportunities

1:05:52 because all the barriers that existed

1:05:55 for you to be able to do this back in the day

1:05:57 have all been torn down.

1:05:59 So to go from somebody

1:06:01 for whom operations meant Salesforce administration,

1:06:05 and then to going from there to saying,

1:06:07 how do I expand the landscape

1:06:09 to knowing how Markov chains work

1:06:12 or knowing how health score modeling works,

1:06:15 to knowing how data infrastructure needs to be set up

1:06:17 to make all of this work,

1:06:20 that transition is like a three to four month focused

1:06:23 learning that you can do at your own pace

1:06:26 based on the needs of your business

1:06:28 and slowly start graduating into it,

1:06:31 get mentorship from other people that are doing it.

1:06:33 It's not impossible, it just needs to happen.

1:06:36 Otherwise to your point, you know,

1:06:38 you would get obsolete if you wouldn't do that.

1:06:42Wrap

1:06:42 - Well, Kishal, this has been an absolute masterclass

1:06:45 and I appreciate you walking through everything,

1:06:48 walking through everything so methodically

1:06:51 and in such detail, going through how to set up

1:06:54 the semantic layer, the context graph, vector database,

1:06:57 tying it all together in the AI applications

1:07:00 that it actually unlocks for the organization,

1:07:03 as well as painting the picture

1:07:05 for what the future of RevOps is.

1:07:07 So Kishal, thank you so much for being on the podcast

1:07:10 and bringing us everything.

1:07:11 Jake, thank you for being my technical backup here,

1:07:13 making me feel way safer

1:07:15 to have this conversation with Kishal.

1:07:17 And I can't wait to see what you and circle do next.

1:07:22 - Thank you so much, it was such a pleasure.

1:07:24 I'm really honored to be a part of this.

1:07:25 I appreciate that.

1:07:28 (whooshing)