The LeanScale Podcast · Episode 99

Why Deflection Is the Wrong Way to Measure AI

Dvir Ginzburg of Encore AI on the metric that tanks revenue, why 'acts human' is the real moat, and the customer who lied to an AI agent

Dvir Ginzburg · Founder & CEO · Encore AI Hosted by Anthony Enrico
Published Updated 00:40:08 28 min read 5590 words
Executive Summary

The one-paragraph brief, extended

Why this conversation matters — and who should spend the hour.

The most popular way to measure a customer-facing AI agent is deflection — how many people it stopped from reaching a human. Dvir Ginzburg's line is that this is like grading a website by how fast people leave it. He is founder and CEO of Encore AI, which builds revenue-generating conversational AI for banks and lenders across voice, chat and IVR, running millions of calls a month; he holds a PhD in geometric deep learning and spent years as a recommendation systems researcher at Microsoft.

His objection is structural: no business has ever measured success by not passing something to a human, and for origination, onboarding, upsell or cross-sell the right metric is satisfaction, revenue or conversion. The travel-insurance case is what deflection costs. That industry measures agents on premium per day; humans were held to it, AI agents were not — deflection rose and PPD dropped sharply, so companies replacing people with agents watched business results decline. Dvir also flags the economics: quoting an investor's line that if you sell a hundred dollars for ninety people will buy, he notes LLM players and wrapper vendors are selling tokens at a loss, and some businesses that chose them are stranded on tech that no longer progresses.

His alternative starts with what already happens in the business. Encore's interaction-mining layer classifies every customer-facing conversation — email, chat, text, voice — as an observability tool first, with the agent stack sitting on those insights. That is how they tell a CEO which flow to build first: not a hunch, but where upsell and cross-sell opportunity is largest. Anthony's framing, which Dvir endorses twice, is that the right proxy is whatever KPI you would put on a human in that role — there is no silver-bullet metric for an employee either, and the first question is what the job is.

On differentiation he is blunt that the foundation model cannot be the moat. Voice quality is commodity, and problems that were hard three months ago — handling interruptions — get solved by an OpenAI release weeks later. What cannot be copied is the client's conversation data. Encore's agents clone top performers: if the best reps tell a joke before the product terms, the agent picks it up; if they have analogies for explaining refinancing to a first-timer, it clones the tactic. Reading terms out like a generic chatbot makes people hang up. He is careful about the ceiling — a nineteen-year VP of collections cannot be beaten, only learned from — and the gains come from reps two months in who skipped training. One retirement fund reported a few million dollars of new funds every hour through the agent, at higher conversion than its human staff.

The signal he most wants is trust: a prospect who said he needed to speak with his wife and hung up rather than say no, exactly as people behave with humans. The real ceiling now is integration rather than capability — core banking systems and obsolete CRMs block the last step. Two further barriers: LLMs treat quantity as quality, so training on raw call archives mostly fails because most conversations are bad; and the channel is shifting to conversational, his own experience of texting Jim his banker rather than opening an app being the model enterprises now want to give everyone. On build versus buy the answer is the 80/20 rule — a great demo is 80%, and the remaining 20% is grunt work like identifying every carrier's voicemail in every language. His closing frame is effectiveness over efficiency; cost reduction, he says, is "truly 2024".

Key Takeaways

15 things worth stealing

The load-bearing ideas, each with the business implication and who should care.

01

Deflection measures the wrong thing entirely

Deflection counts how many customers were prevented from reaching a human. Dvir's analogy is grading a website by how fast people leave it, and his structural objection is that no business has ever measured success by not passing something to a person.

Why it matters: For any revenue-bearing flow — origination, onboarding, upsell, cross-sell — replace deflection with satisfaction, conversion or revenue. Deflection can only tell you a cost story.

Revenue ExecutivesCustomer SuccessRevOps Leaders
02

The travel-insurance case shows deflection tanking revenue

Travel insurance measures agents on premium per day. Human reps were held to PPD; AI agents were not measured on it for roughly three years. Deflection rose and PPD dropped dramatically — companies replacing people with agents watched business results decline.

Why it matters: Where a human role carries a revenue metric, hold the agent replacing it to the same metric from day one. An unmeasured dimension is the one that silently degrades.

Revenue ExecutivesFoundersSales Leaders
03

Measure an agent the way you would measure the human in that seat

There is no silver-bullet metric for an employee either — the first question is what their job is. Anthony's proxy, which Dvir endorses emphatically, is to take the KPI you would put on a human in that role and apply it to the agent.

Why it matters: Start metric design from the role, not from the technology. It also means an agent can carry several objectives at once, as a human would.

RevOps LeadersRevenue ExecutivesCustomer Success
04

The subsidised AI party has an expiry date

Dvir quotes an investor: sell a hundred dollars for ninety and people will buy. LLM players and thin wrapper vendors are selling tokens at a loss, which is not viable — and businesses that chose those vendors are already stranded on tech that does not progress.

Why it matters: Diligence a conversational AI vendor's unit economics, not just its demo. A price below the underlying model cost is a signal about durability, not value.

FoundersRevenue Executives
05

Start with observability, not with the agent

Encore's interaction-mining layer classifies every customer-facing conversation — email, chat, text, voice — before any agent is built, and the agent stack sits on those insights. It is how they tell a CEO which flow to build first: where upsell and cross-sell opportunity is largest, not where the CEO has a hunch.

Why it matters: Dvir's blunt test: if you are building customer-facing Gen AI and have not first understood what already happens in your conversations, you are doing it wrong.

Revenue ExecutivesRevOps LeadersCustomer Success
06

The foundation model cannot be your moat

Text-to-speech quality is commodity, and hard technical problems get solved out from under you — interruption handling was a major challenge three months before recording and an OpenAI release solved it. What Microsoft, Anthropic and OpenAI do not have is your conversation data.

Why it matters: Build the advantage on proprietary interaction data. Anything resting on model capability is a temporary lead measured in months.

FoundersRevenue ExecutivesRevOps Leaders
07

'Sounds human' is solved; 'acts human' is the battleground

Voice models now laugh, pause and carry natural affect, which Dvir describes as slightly freaky even to him. The differentiator has moved to behaviour — whether the agent works the way a good rep works.

Why it matters: Stop evaluating vendors on voice quality, which no longer separates anyone. Evaluate on whether the agent reproduces what works in your business.

Revenue ExecutivesCustomer SuccessSales Leaders
08

Clone the top performers, not the manual

If the best reps tell a joke before sharing product terms, the agent picks it up; if they have specific analogies or a way of explaining refinancing to a first-timer, the agent clones the tactic. Reading terms out in a scripted way makes people hang up saying they cannot deal with a bot.

Why it matters: Training material describes the job; top-performer calls show it being done. AI built from manuals fails on revenue-generating interactions, which is where the difference shows.

Sales LeadersRevenue ExecutivesCustomer Success
09

Agents beat the bottom of the bench, not the top

Dvir will not promise conversion above humans at their best. A VP of collections with nineteen years' experience is, in his words, so good he would pay a debt he does not owe — she cannot be beaten, only learned from. The improvement comes from reps two months in who skipped training.

Why it matters: Model the business case on lifting the distribution's lower half, not on beating the star. And treat top performers as the seed data, which makes retaining them more important, not less.

Sales LeadersRevenue ExecutivesFounders
10

A customer lying to the agent is a trust signal

A prospect told an Encore agent he needed to speak with his wife, then hung up — rather than say no. Dvir's read is that this is precisely how people behave with humans, and he wants every call to reach that level.

Why it matters: Social friction is evidence of a relationship. Without enough trust for someone to soften a refusal, there is not enough trust for meaningful financial operations either.

Customer SuccessSales LeadersRevenue Executives
11

The ceiling is integration, not capability

Complex information gathering, decisioning rules and third-party scoring integrations are all within reach today. What blocks the last step of a journey is core banking systems and obsolete CRMs and ERPs that were never built for it. Dvir adds that Gen AI is the first technology in fifteen years he has seen boards fund integration housekeeping for — a CIO asking for the same integrations to support an analytics tool would previously have been told it was too complex and would take two years.

Why it matters: Scope conversational AI against your integration surface rather than against model capability. The gap is almost never what the agent can understand. It also means there is a window in which long-deferred integration work can actually be resourced; the justification is Gen AI, the payoff is infrastructure that outlasts it.

RevOps LeadersRevenue ExecutivesFounders
12

LLMs see quantity as quality, so raw call archives make models worse

Most conversations in a call centre are bad. Feeding historical conversations in and hoping for the best produces a model worse than an off-the-shelf one. Only recently has it become possible to manipulate that data into input that actually helps.

Why it matters: Volume of historical conversation is not an asset by itself. The curation step — separating what worked from what merely happened — is where the value is created.

RevOps LeadersRevenue Executives
13

The channel is shifting to conversational, and personal banking is the model

Dvir's own experience: as a company with corporate banking he texts Jim, his banker, and does not care what the mobile app looks like. As an individual he gets a digital form. Enterprises are realising they can now give everyone the private-banking experience through omni-channel.

Why it matters: The benchmark is not a better app but a person who already has your context. Anything that forces re-verification and re-explanation is competing against text-your-banker.

Customer SuccessRevenue ExecutivesMarketing Leaders
14

The 80/20 rule separates a demo from production

Eighty percent gets you an impressive demo. The remaining twenty is grunt work — Dvir's example is detecting every telecom carrier's voicemail in every language so the agent does not leave a message and embarrass the bank. It sounds small and it is a dedicated team's problem.

Why it matters: Weight build-versus-buy on the last mile, not the demo. That twenty percent only comes from failures and from having been on the hook for something production-grade.

FoundersRevenue ExecutivesRevOps Leaders
15

Effectiveness over efficiency — cost reduction is 'truly 2024'

The moment Dvir looks for is a CEO saying not only that the goal was met but that they now see business opportunities they had thought unfeasible. Anthony's framing: stop worrying about efficiency and get better than you were before implementing AI.

Why it matters: Build the business case on capability that did not previously exist rather than on headcount saved. The teams doing that, in Anthony's words, leave everyone else in the dust.

FoundersRevenue Executives
Frameworks Discussed

5 named models

Every framework Jimmy names, defined and time-stamped.

Deflection Is the Wrong Metric

01:39

Measuring a customer-facing agent by how many people it prevented from reaching a human, rather than by the outcome the interaction exists to produce — satisfaction, conversion or revenue.

Dvir's analogy is grading a website by how fast people leave it. The structural objection is that no business has ever measured success by not escalating to a person, and the travel-insurance PPD collapse is what it costs in practice.

Interaction Mining

08:40

A layer that classifies and analyses every customer-facing conversation across email, chat, text and voice, used as an observability tool before any agent is built and as the foundation the agent stack sits on.

It answers which flow to build first from evidence rather than from a hunch, by showing where upsell and cross-sell opportunity is largest — and it is how top-performer tactics get identified for cloning.

Sounds Human vs Acts Human

11:22

Voice realism is commodity and improving on someone else's release schedule; behaving like an effective operator, learned from a specific company's own conversations, is the durable differentiator.

Dvir's evidence is that interruption handling went from a major challenge to solved by an OpenAI release in months. The foundation model cannot be the moat because it is shared; the client's data is not.

Cloning Top Performers

13:32

Identifying the highest-performing reps automatically from conversation data, extracting their playbooks — jokes, analogies, phrasing for complex products — and reproducing those tactics in the agent.

The alternative, scripting from manuals, produces an agent people hang up on. The nineteen-year VP of collections is the boundary: top performers are the lighthouse to learn from rather than the target to beat.

The 80/20 Rule of Production

34:22

Eighty percent of the work produces an impressive demo; the remaining twenty is unglamorous edge-case handling that only experience and failure supply.

The canonical example is identifying every telecom carrier's voicemail in every language so the agent does not leave a message. Dvir extends it to the agents themselves — training makes a loan officer, but calls where people say no and hang up are what teaches the job.

Best Quotes

30 lines worth clipping

Pulled verbatim. Copy or share any of them.

“It's kind of like grading a website by how fast people leave it.”
Dvir Ginzburg 00:30
“Never in history we would measure success by not passing it on to a human.”
Dvir Ginzburg 02:13
“The right question should be, what is the satisfaction score? What are we truly wanting to optimize?”
Dvir Ginzburg 02:20
“We were flooded with CXOs from big organizations coming in and saying, "We just need Gen AI. Board forced us, CEO forced us, we just need Gen AI."”
Dvir Ginzburg 03:29
“Deflection increased, but PPD dropped dramatically.”
Dvir Ginzburg 05:14
“I am replacing, quote unquote, people with agents, but my business results are actually decreasing or taking a hit.”
Dvir Ginzburg 05:20
“One of our investors say that if you sell 100 bucks for 90, people will buy.”
Dvir Ginzburg 06:28
“We are all enjoying the subsidized AI party right now, where we can run as many agents as we want, and it's not too cost prohibitive. But there's going to be a time where that's not the case.”
Anthony Enrico 06:05
“If you are an enterprise that is looking to build Gen AI agents for your customer facing processes and you didn't go through the step I'm just going to describe, you are doing something wrong.”
Dvir Ginzburg 08:18
“We first create an observability tool on what is going on within the business. And then our entire agenda stack sits on these insights.”
Dvir Ginzburg 08:53
“We say that not because of a hunch. It's because we saw that the opportunity, the absence and crosses opportunity there are the biggest.”
Dvir Ginzburg 09:14
“You wouldn't say, hey, there's a silver bullet for how you should measure your employees. It's like, well, first, your next question would be, well, what's their job?”
Anthony Enrico 09:47
“The foundational model cannot be your mode. What can be your mode is your client's data.”
Dvir Ginzburg 12:50
“If their top reps tell jokes before sharing the product terms, our agent will be able to pick up on that.”
Dvir Ginzburg 13:32
“You need to understand what's working, how top performers explain, what generates that connection, and utilize that.”
Dvir Ginzburg 14:40
“It literally lied to the AI agent because he didn't feel comfortable telling it "no". Just as they would behave with a human.”
Dvir Ginzburg 16:51
“If I won't be able to create that level of relationship or trust, I won't be able to have their trust to do meaningful operations with what we provide them.”
Dvir Ginzburg 17:42
“I will not be able to beat her. No AI agent will be able to beat her because she has done it her entire career.”
Dvir Ginzburg 19:20
“They're generating a few million dollars of new funds coming in every hour from our agent.”
Dvir Ginzburg 20:17
“These are the ones we will never be able to replace. We will be able to learn from them. They will be our lighthouse in a sense.”
Dvir Ginzburg 21:36
“Currently, the ceiling lies within the integration layer and not the technical layer.”
Dvir Ginzburg 22:19
“With Gen. AI, it's the first time where you see boards and CEOs move things like bulldozers.”
Dvir Ginzburg 23:36
“LLMs see quantity as quality, which means that if you want to train LLMs on your historical data, it was almost mission impossible.”
Dvir Ginzburg 25:20
“When I have a question as Encore, I just text Jim. This is what I do. I just text Jim, and he replies, and I don't care how the mobile banking app looks like.”
Dvir Ginzburg 28:56
“Enterprises understand that they can provide personal banking experience through omni-channel.”
Dvir Ginzburg 29:45
“You can create an amazing demo with 80%. But then the 20% of grunting comes in.”
Dvir Ginzburg 34:29
“The demo sounds so good, and then you hit production and reward issues arise.”
Dvir Ginzburg 34:56
“AI that simply replicate manuals and training sessions is AI that will fail when it comes to revenue generating interactions because revenue generating interactions are hard.”
Dvir Ginzburg 36:34
“We are not only looking to replace or cost reduction. That's truly 2024.”
Dvir Ginzburg 39:09
“Let's not worry as much about efficiency and let's focus on effectiveness. Let's get better than we were before implementing AI.”
Anthony Enrico 39:28
Practical Advice

What should you actually do?

The playbook, split by the seat you sit in.

Revenue Executives

  • Replace deflection with the outcome metric for the flow — satisfaction, conversion or revenue. Deflection can only tell a cost story.
  • Hold an agent to whatever KPI the human in that seat carried, from day one. The travel-insurance PPD collapse happened because nobody did.
  • Build the case on capability that did not exist before rather than on headcount saved; cost reduction is the 2024 argument.
  • Scope against your integration surface, not against model capability — core banking systems and old CRMs are what actually block the last step.
  • Diligence vendor unit economics. A price below the underlying model cost signals fragility, not efficiency.

RevOps Leaders

  • Stand up conversation observability before building any agent — classify email, chat, text and voice to find where revenue actually leaks.
  • Choose the first flow from where upsell and cross-sell opportunity is largest, not from where the loudest stakeholder is.
  • Do not feed raw call archives to a model. Most conversations are bad, LLMs treat quantity as quality, and the result is worse than off the shelf.
  • Use the current funding window: Gen AI is the first justification in fifteen years that gets a year of integration housekeeping approved.

Sales Leaders

  • Build agents from top-performer calls rather than from training manuals — the manual describes the job, the calls show it being done.
  • Model the business case on lifting the bottom of the bench. The nineteen-year veteran is the seed data, not the target.
  • Watch for social friction as a trust signal: a customer softening a refusal rather than saying no is behaving as they would with a person.
  • Stop evaluating vendors on voice realism. It no longer separates anyone.

Customer Success

  • Benchmark the experience against texting a banker who already has your context, not against a better app.
  • Treat every re-verification and re-explanation as friction the conversational alternative does not have.
  • Keep the AI disclosure — Encore always states the caller is speaking to an agent, and connection still forms within the conversation.
AI Takeaways

How AI actually changes GTM

LeanScale's signature read on the AI-in-GTM question this episode wrestles with.

The thesis

The whole conversation is an argument that AI value is being measured backwards. Deflection optimises for absence of contact, which is only ever a cost story, and the travel-insurance case shows it actively destroying revenue while the dashboard improves. Dvir's replacement is that the moat is not the model but the enterprise's own conversation data — and that the binding constraint on deployment is integration, not intelligence.

Agent & automation ideas

  • An interaction-mining layer classifying every email, chat, text and call to locate where revenue leaks before any agent is scoped.
  • A top-performer extraction agent that identifies the best reps from outcomes and derives their reusable tactics.
  • Proactive outbound voice for events that matter — a significant account change — with per-carrier voicemail detection so nothing is left on a machine.
  • A wingman assistant surfacing top-performer tactics to human reps live rather than replacing them.
  • A conversation-curation step that filters historical calls to the ones worth learning from before any training run.
Operations Takeaways

By function

The same conversation, filtered for RevOps, pipeline/marketing ops, and customer ops.

Revenue Operations

  • .
  • .
  • .
  • .
  • .

Pipeline & Marketing Ops

  • .
  • .
  • .

Customer Operations

  • .
  • .
  • .
Metrics Mentioned

The numbers, with context

Millions of calls per month
Encore call volume

Across voice, chat and IVR, deployed either as fully autonomous agents or as a live wingman to human teams.

40+ globally
Enterprises served

The base of production experience Dvir cites as the source of the product insights and scars that make the last mile tractable.

Deflection up, PPD down sharply
Travel insurance outcome

Premium per day is the industry metric for human agents; AI agents were not held to it for roughly three years, and business results declined.

A few million dollars of new funds per hour
Retirement fund result

Reported by the CEO of a retirement fund that is not among the largest, with conversion higher than the human staff achieved.

$30M raise
Encore funding

Mentioned in passing as the reason the company has corporate banking, in the anecdote about texting his banker rather than using the app.

19 years
Tenure of the top collections performer

The VP of collections Dvir says no agent will beat — the boundary case for what cloning can and cannot reach.

Entities

Companies, people & tools mentioned

Auto-extracted and linked into the knowledge graph.

Companies

People

Tools & software

ChatGPTAI Assistant

Used as the negative example of scripted delivery — reading product terms out the way a generic chatbot would makes customers hang up saying they cannot deal with a bot.

SlackTeam Messaging

Anthony's answer to how he would contact a friend, in the exchange establishing that the natural channel for a question is a message rather than an app or a form.

Methodologies referenced Interaction mining before agent designTop-performer cloningRole-equivalent agent KPIs
Frequently Asked Questions

Straight answers

Generated from the conversation, marked up for search and AI extraction.

Why is deflection the wrong metric for customer-facing AI?

Deflection counts how many customers were stopped from reaching a human, which Dvir Ginzburg compares to grading a website by how fast people leave it. No business has ever measured success by not escalating to a person. For any flow that exists to produce an outcome — origination, onboarding, upsell, cross-sell — the right measures are satisfaction, conversion or revenue. Deflection became popular because boards demanded Gen AI quickly and cost reduction is the easiest thing to prove.

What happens when you optimise an AI agent for deflection?

The travel-insurance case is the clearest example. That industry measures agents on premium per day — how much is upsold and cross-sold onto a basic policy. Human reps were held to PPD, but for roughly three years AI agents were not. Deflection rose while PPD dropped dramatically, so companies replacing people with agents watched their actual business results decline. The general lesson is that where a human role carried a revenue metric, the agent replacing it needs the same metric from day one, or the unmeasured dimension is the one that silently degrades.

How should you measure an AI agent instead?

Take whatever KPI you would put on a human doing that job. There is no single silver-bullet metric for an employee either — the first question is always what the role actually is, whether sales, operations or something else — and the same reasoning applies to agents. This also means an agent can hold several objectives at once, in the same way a bank might incentivise staff to promote a new payments product alongside their existing targets.

What is interaction mining?

A layer that classifies and analyses every customer-facing conversation in an organisation — emails, chats, text messages and voice calls — to create an observability picture of what is already happening before any AI agent is built. It serves two purposes: identifying which flow to automate first based on where upsell and cross-sell opportunity is largest rather than on a hunch, and identifying how top performers actually win so the agent can reproduce their tactics.

Why can't the foundation model be your competitive moat?

Because it is shared and it improves on someone else's schedule. Text-to-speech quality is already commodity, and hard technical problems get solved out from under you — handling interruptions was a significant challenge three months before this recording and an OpenAI release solved it. What Microsoft, Anthropic and OpenAI do not have is your organisation's conversation data. Companies that cannot use their own interaction history to differentiate will not progress.

Can AI agents outperform human sales or collections staff?

Against the top of the team, no. Dvir describes a VP of collections with nineteen years' experience whose calls are so effective he says he would pay a debt he does not owe — no agent will beat her, and she functions as a lighthouse to learn from rather than a target. The gains come from the rest of the distribution: reps two months into the job who skipped training. One retirement fund reported generating a few million dollars of new funds every hour through the agent, with conversion higher than its human staff overall.

What is actually limiting AI agents today?

Integration, not capability. Complex information gathering, decisioning rules and third-party scoring integrations are all achievable now. What blocks completing a journey end to end is core banking systems and obsolete CRMs and ERPs that were never built to support it. Notably, Gen AI is the first justification in more than fifteen years that gets boards to fund a year of integration housekeeping, where the same request for an analytics or observability tool would previously have been rejected as too complex.

Why doesn't training an AI on your historical call recordings work?

Because LLMs treat quantity as quality, and most conversations in a call centre are not good conversations. Feeding the archive in wholesale produces a model that performs worse than an off-the-shelf one. Only recently has it become possible to manipulate historical conversation data into input that genuinely improves a model — which means the curation step, separating what worked from what merely happened, is where the value is created rather than in the volume itself.

What is the 80/20 rule of taking conversational AI to production?

Eighty percent of the work produces an impressive demo; the remaining twenty percent is unglamorous edge-case handling. Dvir's example is voicemail: for proactive outbound calling, every telecom carrier and every language has a different voicemail system, and the agent needs to identify each one so it does not leave a message and embarrass the bank. It sounds trivial and it is a dedicated team's problem. That last mile only comes from failures and from having been on the hook for something production-grade, which is the core of the build-versus-buy calculation.

Full Transcript

The whole conversation

Broken into chapters, searchable, verbatim from the audio. Speakers inferred (not diarized).

00:00Cold open + intro

0:00 Probably in a few months from now, when people will listen in, the tech would be even more advanced.

0:30 It's kind of like grading a website by how fast people leave it.

0:34 You know, the most popular method to measure agents today, customer-facing agents, is deflection.

0:41 You would go to any of the most popular solutions.

0:44 The first method they would be proud of is what is our deflection.

0:48 So in this space you're talking about, anyone can make a voice sound human-ness.

0:54 What's the difference between sounds human and acts human?

0:59 You need to understand what's working, how to performance explain, what generates that connection, and utilize that.

1:07 And that's something that not you, not Microsoft, Antropic, OpenAI, no one have.

1:14 They have it.

1:15 And the question is, how can you utilize that?

01:22The deflection trap: grading AI by how fast people leave

1:22 Daveer, you told me that measuring AI by how many people it stops from reaching a human is kind of like grading a website by how fast people leave it.

1:32 I'd love if you could unpack that for us, and why does it wind you up?

1:37 Of course, yes. Thank you very much, Anthony.

1:39 So, you know, the most popular metric to measure agents today, customer-facing agents, is deflection.

1:47 You would go to any of the most popular solutions, the first metric they would be proud of is what is our deflection rate.

1:55 Meaning how many clients are not being moved to human reps, and the inquiry is finished with speaking with a bot.

2:03 Now, I think it makes sense now that we spoke about it, that measuring that as a success of an AI agent doesn't make sense.

2:13 Never in history we would measure success by not passing it on to a human.

2:20 The right question should be, what is the satisfaction score?

2:25 What are we truly wanting to optimize?

2:29 Because if we are speaking on an origination portal, onboarding, buying a new product, or upselling or cross-selling a client, well, deflection shouldn't be the metric, right?

2:42 It should be the satisfaction score, or even more important, the revenue or conversion rate we were after.

2:51 Yeah, and I think you would compare it to similar metrics you'd have in a human setting, so I fully agree you wouldn't measure anybody on your team that way.

3:00 When do you think the market for this kind of flipped, and people are now going to be expecting more than just deflection away from engaging your team to,

3:10 "Hey, we want these agents tied to actual revenue metrics"?

03:14From 'show my board I have AI' to 'what's the ROI?'

3:14 Yeah, amazing. So, at first, and sounds old, right, but two years ago when Gen AI started to become stable and voice agents became a thing, everyone wanted it.

3:29 And we were flooded with CXOs from big organizations coming in and saying, "We just need Gen AI. Board forced us, CEO forced us, we just need Gen AI."

3:40 And that pushed them to find for very basic measurements to quantify by.

3:48 So deflection was a very fast and immediate aspect because we want to show cost reduction.

3:53 Cost reduction is always easier to prove out.

3:58 Well, it's a bit harder than that, but I believe that we will touch on that in a moment as well because AI costs money, right, at the end.

4:05 But at first, they would measure for deflection, and this is what happened for the last 24 years.

4:13 Now, luckily enough, the ecosystem is maturing.

4:20 So, you know, eventually, although the US is a very big country, we see same vendors and same prospects coming into the same venues.

4:32 And the CEOs that at first just said, "Well, I want AI to show my board I have AI," are suddenly asking, "Well, what is the return on investment? What is the value I'm going to get?"

04:46Travel insurance: deflection up, revenue down

4:46 And just as you said, let me give an anecdote. When discussing travel insurance, one of the most popular metrics to measure agents by is PPD, Premium Per Day, right?

4:58 How much you were able to upsell and cross sell on the basic travel insurance policy.

5:04 Now, somehow, for the past three years, no one was measuring AI agents by that metric, although the humans were being measured by that metric.

5:14 So, you know what happened? Deflection increased, but PPD dropped dramatically.

5:20 So, suddenly, businesses said, "Wait, I am replacing, quote unquote, people with agents, but my business results are actually decreasing or taking a hit."

5:32 So, today, the aim is much more business oriented than just our CMO likes to say, "Sprinkle Gen AI over existing tools."

5:45 It's actually doing something meaningful that brings business value back to the organization.

5:52 Yeah, and if deflection isn't the right scoreboard, how exactly should you be measuring them?

6:00 And I think one thing you brought up that's important in this context too is the cost is real.

6:05 And I think we're all enjoying the subsidized AI party right now, where we can run as many agents as we want, and it's not too cost prohibitive.

6:15 But there's going to be a time where that's not the case.

6:18 But how should you be measuring the effectiveness? What exact KPIs should you be looking at every day to make sure that your agents are actually built to do the things you need them to do?

06:28The $100-for-$90 problem and the subsidized AI party

6:28 Yeah, amazing. So, first of all, I have a funny anecdote. One of our investors say that if you sell 100 bucks for 90, people will buy.

6:38 And I think that we are in that point today where startups and the big LLM players, they are basically selling, they are selling with a loss to themselves tokens.

6:54 So, of course, that if you are a chatbot vendor and you are basically an LLM wrapper and you are cheaper than simply using the LLM itself, clients will come.

7:06 But it simply doesn't make sense. It's not viable. And many businesses that shows those players are now left hanging with tech that doesn't do anything or it doesn't progress.

7:17 So I agree with you 100% there. In terms of measurement, I think this is another maturity step that the ecosystem went through because today I don't have a silver bullet.

7:32 I can tell you, yes, this is the measurement. It's not as easy as deflection. But what is easier is that you can and need to sit with the business.

07:43Deployment strategists and interaction mining

7:43 And sometimes you will even have a dedicated role to it. Deployment strategist.

7:48 And deployment strategist is the person that comes into the business to the organization and understands together with them what are we trying to optimize in each and every flow.

8:02 Now, this is not being done in a manual way like a consultancy firm. No.

8:09 We and now I am speaking about what we do at Encore. I hope that others are doing it differently. But again, and of course, I'm not objective.

8:18 If you are an enterprise that is looking to build Gen AI agents for your customer facing processes and you didn't go through the step I'm just going to describe, you are doing something wrong.

8:33 And that thing is understanding what is going in the business already.

8:40 We have built a layer called interaction mining, which is our ability to classify and see through all of the customer facing conversations within the organization.

8:53 It doesn't matter whether it's emails, chats, text messages or voice calls. We analyze them all. We first create an observability tool on what is going on within the business.

9:07 And then our entire agenda stack sits on these insights.

9:14 So when we are coming to the CEO telling him this should be the first flow to build together, we say that not because of a hunch, not because of the CEO really doesn't like that the lending staff.

9:28 It's because we saw that the opportunity, the absence and crosses opportunity there are the biggest.

9:37 Yeah. And I think maybe a good proxy is taking a look at the KPIs you put on a human in that given role or who's responsible for that given process.

9:47 You wouldn't say, hey, there's a silver bullet for how you should measure your employees.

9:50 It's like, well, first, your next question would be, well, what's their job? Are they in sales? Are they in operations? Are they developer?

9:57 Then you can start to have the conversation of what those KPIs should be. And it's probably the closest we can get to how should you measure if your agents are being effective?

10:07 One hundred percent. One hundred percent. And by the way, I think that that's an amazing point because that can change is really right.

10:15 We are working again, but we are working with a bank. They just opened a new product line, a payment solution for for other vendors that on board their platform.

10:26 And they wanted more engagement there. So they incentivized their employees to start promoting that payment solution.

10:37 Now, if they are now incentivizing their humans to do that, there are ways to incentivize agents as well.

10:46 So suddenly the agent doesn't need to only have one metric. It is optimizing. It can have multiple ones. And yes, as you said, it's business analytics work that is required.

10:59 It's mandatory to play the long game of Gen AI in customer facing interactions.

11:07Sounds human vs. acts human — your moat is client data

11:07 So in this space, you're talking about anyone can make a voice sound human now that that's pretty widely available.

11:17 What's the difference between sounds human and acts human?

11:22 Yeah, spot on spot on because A, enterprises have matured enough to understand that point solutions within this space are becoming irrelevant.

11:38 And if all that you do is a voice, I want the bigger players and the vertical players are going to eat you up.

11:50 And second, because the tech is advancing so fast, having a good text to speech model is irrelevant now.

12:01 Again, I don't want to be too tied to the current times. Probably in a few months from now, when people will listen in, the tech will be even more advanced.

12:11 But some of the technical challenges that we faced three months ago, and I don't want to be too technical, but things like interruptions, right?

12:20 You are starting to speak, suddenly the agent speaks, you want to run over it. Sounds easy. Three months ago, that was a huge challenge.

12:28 How are you being able to manipulate the agent to stop speaking, et cetera, et cetera?

12:33 OpenAI came a few weeks ago with a new model and that's been completely solved, right? So you, just like with more rule with chips, you are finding yourself solving for the next problem that big LLMs are going to solve in a moment.

12:50 And so the foundational model cannot be your mode. What can be your mode is your client's data and companies that won't be able to utilize the client's data will not be able to progress.

13:07 And what we focus on at Angkor is understanding what's working within the enterprise today because not only because I love my partners and clients, but because I know that if tomorrow a new vendor will come in and we'll knock on their doors and tell them, "Oh my God, I have that amazing tech. You should check it out. Here is a demo I created for you."

13:32Cloning your top performers

13:32 That demo is going to be non-personalized and not fitted for their culture and need. The agent they already have today is literally cloning their top performers because the agents that we have built together, if their top reps tell jokes before sharing the product terms, our agent will be able to pick up on that.

13:56 If they have specific analogies or a market brand that comes with each of their products, our agents will be able to clone that tactic.

14:08 If they have a specific phrase or a way of explaining a complex request like, "We are working a lot in the refinancing space."

14:18 And if you are a banker, refinancing is so easy. If you are an individual, it might be the first time you are refinancing a loan ever. That's not that easy.

14:30 If we will just, in a scripted way, tell the agent what to say when a client is asking, "Wait, I don't understand the terms. How does it even work?"

14:40 If you will just read out, like Che GPT, clients will hang up. They will say, "I can't deal with this bot." So you need to understand what's working, how top performers explain, what generates that connection, and utilize that.

14:56 And that's something that not you, not Microsoft, Antropic, OpenAI, no one has. They have it. And the question is, "How can you utilize that?"

15:07The customer who lied to the AI about his wife

15:07 And how human can this actually get? Because I imagine there's going to be a lot of hesitation to how good one of these agents could be.

15:19 And does the person on the other side really connect with an agent the same way that they would connect with a human being?

15:27 It is becoming a bit freaky even for me, to be honest. Just lately, all big competitors for text-to-speech models, release models that can laugh, pause, moan.

15:40 And that really created that feeling that you are speaking with a human because suddenly you are telling a joke and it laughs in a way that is very natural.

15:50 So that was a "ha" moment. I will say that from compliance perspective, we always share that they are speaking with an agent, an AI agent.

16:04 But because we have that interaction writing piece and because we are cracking the conversation with a joke, sometimes speaking about sports,

16:14 that connection starts very fast within the conversation. That's A. And B, we have so many stories, let me give you one,

16:26 where we are, as I share, we are focusing a lot around deposits and retirement funds and loans.

16:35 And one of our AI agents gave a product proposal to a prospect and the client said something like, "Well, I need to speak with my wife, hang out a second, and then hang up."

16:51 So it literally lied to the AI agent because he didn't feel comfortable telling it "no". Just as they would behave with a human.

17:01 And then, of course, we cannot listen to any of the calls. Of course, we currently run millions of calls a month.

17:11 I would love to hear some of them. But sometimes our clients, under a very specific NDA within their environment, everything is contained, are sharing some.

17:25 And it helps us to improve so much because we understand what is working and what doesn't.

17:35 And we understand because I want every call to be like that. I want that if someone wants to lie to my AI agent, it will be like that.

17:42 Not saying no, I want them to say I need to speak with my wife and then hang up. Because if I won't be able to create that level of relationship or trust, I won't be able to have their trust to do meaningful operations with what we provide them.

18:04 I think that's probably one of the biggest barometers when a human feels like it has to lie to the AI. I think that's a good sign that it's operating in a pretty human way.

18:13 So the connection is starting to get there. The models are catching up where they're putting the things in there that are important and relevant. You have the context of the company data.

18:24Do agents actually beat humans?

18:24 How do these actually perform? I think when you're measuring it, let's call it conversion, upsell, cross-sell, compared to the humans, how can these agents perform? Are they on par, better or worse? Where do they land?

18:41 We would never promise conversion rates that are better than humans. Definitely not if you max them out. What I mean by that is that eventually what we do is that we come to the business and we identify the top performers automatically, we understand their playbooks and then we try to resemble them.

19:06 So there will always be those top performers, the people that excel. Sometimes these are the people we work with, the debt collector, the VP of collections, she's there for 19 years.

19:20 I listen to some of her calls. She's doing her job so well. I don't have a debt with them, but I will pay. Truly, like she's so good. That's impressive. Now, I will not be able to beat her. No AI agent will be able to beat her because she has done it her entire career.

19:41 But the ones that are doing that for two months and are about to leave, the ones that skip training and are not doing that because everyone in the debt collection agency are so packed with work so they need somewhere to get on calls.

19:58 These we are able to improve upon. And I just had a call with the CEO of a very big retirement fund. They shared with me that we are now in a rate of generating a few million dollars of new funds coming in every hour with our agent.

20:17 And they are not a big retirement fund. Definitely not one of the big 10. And still, they're generating a few million dollars of new funds coming in every hour from our agent.

20:29 That was probably the first time when a CEO shared with me not only that I was able to scale and reach people that I wasn't able to scale before.

20:41 I am able to generate more revenue and more conversion rate with the agents than with my human staff.

20:51 So in that case, the agents are outperforming the human team.

20:55 If you look at the human team as a group, as a whole, yes, they have the potential of doing that.

21:03 You will always have those individuals that excel. They're doing it for years. They have their playbooks.

21:12 They know to identify things that even their colleagues, they can't identify. It just happened. You know, sometimes it's like reading a book.

21:21 You don't know what makes it such a good book, but it just works, right?

21:27 This is exactly what happens with these top performers. You listen to their calls and you say something magical happened here.

21:36 I don't know how to imitate it. And these are the ones we will never be able to replace. We will be able to learn from them.

21:44 They will be our lighthouse in a sense.

21:49 Yeah, you still need that seed data to model off of, and I think that makes sense as kind of like the limiting factor of success.

21:58 How complex can these agents go? There's clearly some really good, I'll call them transactional level conversations.

22:08 It's maybe one conversation at a time, simple products at how big of a ceiling do these agents have in terms of being able to handle more complex tasks?

22:19The real ceiling is integration, not technology

22:19 Yeah, yeah, amazing. So, currently, the ceiling lies within the integration layer and not the technical layer.

22:31 Meaning most of the times where we are stuck away from doing more, it's because the core banking systems or the obsolete CRMs and ERPs aren't allowing us to do that extra step to complete the journey end to end.

22:56 But when it comes to the complexity of gathering information, having complex decisioning factors, having rule base to decide what to ask and what not to ask, all the way to integrating to third party scoring systems to decide what to offer and what not to offer, all of that is within the reach of AI agents even today.

23:24 The only gap that we have is the technical aspects of reaching out to integrations that the old generation haven't sold for.

23:36 And by the way, I will say another thing that I'm constantly amazed by, and I'm selling tech in various ways and systems for more than a decade and a half now.

23:48 And with Gen. AI, it's the first time where you see boards and CEOs move things like bulldozers.

23:59 So they would see an objective of "I want this Gen. AI to do X" and they will do whatever it takes to make sure it has the necessary integrations.

24:11 And let me explain what I mean by that.

24:14 Five years ago, if the CIO would come and say "I have that amazing observability tool" or "I have that amazing predictive analytics or recommendation engine tool, but I need integrations to 1-2-3", the board would say "too complex, I can't have it, it's too expensive, it will take us two years".

24:35 Now, when the CIO is coming and saying "I have a new Gen. AI tool, but we need to do housekeeping and house cleaning for a year to make sure that we are ready for that", people are investing in these digital transformations to go through Gen. AI transformation.

24:55 And that's a very unique moment that I haven't experienced for a long time.

25:02 What is holding things back now where before maybe it was, hey, laughter, certain nuances of human language, the pausing, are there any big barriers now that are holding anything back that you think might be lifted?

25:20The next barrier: 'LLMs see quantity as quality'

25:20 Yeah, yeah, yeah, sure. So, two things there. First of all, it was so hard to utilize internal data only a few years ago because, and it might sound complex, so bear with me, LLMs see quantity as quality, which means that if you want to train LLMs on your historical data, and this is exactly what we have patterns over,

25:49 it was almost mission impossible because even with humans, most of your conversations are not good conversations. If you would look at conversations within call centers, most of the conversations are banned.

26:03 So, just feeding in these conversations and hoping for the best actually creates LLMs that are worse than simply having off-the-shelf models. So, only now we are in a point where we can take existing data and manipulate it in a way that actually serves as good input for the model than have that generate the aggregation in result.

26:29 So, I think that's the first thing that is being lifted. How can you create something that is not only, doesn't only sound well, and is not only local in terms of multilingual and all that, but can actually generate that personality that you are after.

26:48The omni-channel shift: 'text Jim'

26:48 The second thing that I will say is that up until a few years ago, it was the consensus that you communicate with the enterprise, whether it's a bank, insurer, whatever it may be, through your mobile banking app, through the website, or you would call.

27:12 And today, the concept of omni-channel or going fully conversational over text messages, maybe WhatsApp or Telegram, that's something that we see the younger generation and younger neo-banks are doing that is going to become more and more popular.

27:32 So, imagine, let me ask you that, when you are trying to check out your balance in your bank, how do you do that?

27:44 I have an alert that comes every morning. So, I get the text message. I like it to be text because my emails are flooded.

27:51 And if it's the middle of the day and there's something specific I'm looking for, then I'll open up my app or log in if I'm working and on my laptop.

28:00 Exactly. But if you need to check something with your friend, how would you do that?

28:08 If I needed to check on something with a friend, you are having a thought, you need to check something with your friend, you would probably text them, right?

28:21 Yeah, text them or Slack if a lot of my friends I work with. So, you just shoot them Slack. Exactly. You just shoot a message over Slack.

28:31 A shift that we are now seeing, and by the way, you know, at Encore, I have my own personal bank account, of course, I don't have private banking, as most of the population don't.

28:44 But as Encore, we just had a very big raise, we raised 30 mil. We are part of the corporate banking of one of the big banks, so we do have a banker.

28:56 Now, when I have a question as Encore, I just text Jim. This is what I do. I just text Jim, and he replies, and I don't care how the mobile banking app looks like.

29:12 I just get answered, and if I need to open a deposit account for Encore, I just speak with Jim, and he gets that sorted out, and if Jim needs to, he calls me, and he goes over the process with me over the phone until I get it.

29:29 And I can ask a million questions, and he will just sort it out for me. As an individual, this is the digital form, fill it, and you are done. You can't, your problem.

29:45 The shift that we are now seeing is that enterprises understand that they can provide personal banking experience through omni-channel.

29:58 I think that's tremendous. Same, they have private banking on the business side, and there are things where it's just nice to text someone who already has all the context, already knows everything.

30:11 I don't need to log into the account. I don't need to re-explain anything or go through extra verifications or all the blah, blah, blah.

30:17 So I think having that type of experience, absolutely, I think I would be looking for as a consumer, and if it's something that can be built at scale, I think that adds tremendous value, and it's exciting to see some of these capabilities open up where people can start to stitch these things together.

30:35Build vs. buy and the 80-20 rule of production

30:35 But what it leads me to question, and I would love your perspective on this, if a lot of the core models have gotten really, really far along, solved many of the problems of communication, especially in the omni-channel environment, what is the build versus buy equation?

30:57 Because before, for many things, it was extremely prohibitive to build things. You need to hire your own developers, engineers.

31:07 Now, at LeanScale, we built plenty of internal applications and tools, even multi-tenant applications for our customers, things that they can start using without having a real full dev team.

31:23 In this context, though, what's the build versus buy decision when you're evaluating something like conversational AI?

31:32 Utilizing your own conversations to train the models is still an extremely complex task, and I think there is a very good reason why most still optimize deflection and why almost no one is using these existing conversations to build their stack.

31:51 So I think that's, first of all, but I'm not naive.

31:56 People are now listening to this podcast, people go to conferences, people understand this is the future, and soon enough, that will be hopeful.

32:05 I'm truly hopeful because it doesn't matter whether we are first, second, or tenth.

32:12 The ecosystem will understand eventually, it might be in a month from now or 12 months, but that without optimizing on your existing interactions, it's just mute, it's irrelevant.

32:25 So I would say that that's first in the build versus buy dynamics.

32:33 Startups and companies will always be more innovative in the capabilities that NLMs can allow you.

32:41 The second in the build versus buy is that, like with everything, Gen. AI is not a magic bullet to good product.

32:50 And we are now working with more than 40 enterprises globally.

32:56 The things that we have learned, and the scars that we have, and the product insights that we own, you know, even the funny things, you are having an outbound call.

33:10 Because you want to have proactive, we are very strong in proactive communication.

33:16 So you just mentioned a text on balance, but something changes your account and your product line, a significant drop.

33:26 We are really good in sending personalized calls to inform people that prefer to have it in a call.

33:33 And of course, proactive communication via voice was simply irrelevant out of reach for almost all banks three years ago.

33:43 Because no one would put the resources to call half a million clients to tell them that the bank just launched a new product, right?

33:52 It's very unfeasible.

33:53 Suddenly with Gen. AI, it is. But for each telecommunication vendor, Verizon, AT&T, whatever it may be, and for each language, you have a different voicemail.

34:07 And if you don't want to spam the voicemail, you need to have a very specific identifier.

34:13 We just reached voicemail. I don't want my AI to speak with a voicemail and embarrass us or the bank, right?

34:22 It sounds so small, but this is exactly the 80/20 rule of software.

34:29 You can create an amazing demo with 80%.

34:34 But then the 20% of grunting comes in and you understand that you are the CIO of a bank and you need to have a dedicated team to identify different voicemails and how to handle each one.

34:48 Well, we solved it. And again, I'm not speaking.

34:56 It's not part of the equation today. It's just a build versus buy. The demo sounds so good, and then you hit production and reward issues arise.

35:05The last mile + wrap: effectiveness over efficiency

35:05 And I think that's where a lot of people really don't realize just like everything.

35:12 It's that last mile that is so difficult that only comes through failures and experiences and being on the hook for having something production grade.

35:23 And yes, it's easy to show a demo. It's easy to kind of spin up an idea of something.

35:31 But like you mentioned, getting all of the edge cases in the gotchas is very difficult to do.

35:37 So I think that makes sense in this case. Like, OK, if one, if you have a level of seriousness with your product and maybe compliance you've got to follow and things like that, like there are certain things that you cannot mess up.

35:52 Then you don't want to vibe code your way to that. Or if the volume and scale is just at a level that's unreasonable to manage on your own, I think that makes a ton of sense.

36:03 And by the way, I will tie it to what we do at Anchor because just like doing the 80/20 from demo to production, you can do all the training necessary to become a loan officer.

36:21 But then you hit the ground running and you start making calls and people say no and people hang up and you start to understand what works and what doesn't.

36:34 And AI that simply replicate manuals and training sessions is AI that will fail when it comes to revenue generating interactions because revenue generating interactions are hard.

36:49 So just like that, we bridged the 80/20 for LLMs going from a demo prompt to a production that mimics or clones what really works within the business.

37:07 Daveer, I think this is a huge topic, especially looking into build versus buy the level of effort it takes to get it that last mile. And I appreciate all of the nuances going.

37:20 The difference between sounding human and acting human and not just mimicking a person on your team, but mimicking the best people on your team and using that seed data to make something that actually makes a huge difference for the customer.

37:34 And I'm also excited that there's an opportunity to start to democratize some of these higher and higher level experiences that are usually reserved for only a handful of people, but now AI can really bring that to everyone within a given organization.

37:50 So Daveer, thank you for all of the knowledge and AI transfer on this one. And I'm really excited to see congrats on the fundraise as well. That's very, very well done. And I'm really excited to see what you all do next with your funding round and as things just change at neck break speed.

38:13 Amazing. Yeah, I had a really fun time. Thank you very much. Yeah, I think we are truly in a singular moment in time where everyone understand the impact.

38:27 And now tech converges well into a place where you can execute your dreams. And maybe just to define a sentence, I will say that the thing that excites me the most is that we have that moment where we go live with one of our clients, banks, lenders, whatever it may be.

38:47 And suddenly I go into the CEO's office and he says, not only he or she or she say, not only that I was able to achieve the goals I have set to myself, but I now understand that I have business opportunities I thought are not realistic or unfeasible before.

39:09 And that's what we are looking for. We are not only looking to replace or cost reduction. That's truly 2024. It's how can I give the CXOs the atmosphere and feeling that they can do more than what they ever dreamt is possible.

39:28 I think that's the most important distinction. Let's not worry as much about efficiency and let's focus on effectiveness. Let's get better than we were before implementing AI. Let's do more for our customers. Let's add more value.

39:45 And the teams that are doing that are going to leave everyone else in the dust.

39:52 100%.

39:54 Daveer, thank you so much and can't wait to see what you all do with your next stage of growth.

40:01 Perfect. Thank you very much. It was a pleasure.