37:26 - Yeah, totally.
37:28 And you know, the way that you would even go
37:30 about developing this is that you never wanna embark
37:32 on an eight month project
37:33 where you're gonna do all of this at once.
37:35 Because by the time month number six arrives,
37:37 like world level have changed.
37:39 Some tech would have gone obsolete,
37:41 some new stuff would have arrived
37:42 that makes some of the stuff you did
37:44 in month number three useless or onerous or whatever.
37:48 Instead, what you wanna do is have this long-term vision
37:51 where you know these are the high level pieces
37:54 you're gonna need, have that vision
37:56 really, really strongly developed.
37:58 But then as you get to the individual steps,
38:00 look at what are your immediate needs and priorities
38:03 and where do they fit in that vision, right?
38:06 So if today even your numeric data is a mess,
38:09 I wouldn't worry about vector databases or context graphs,
38:13 I would clean up the numeric data right away.
38:15 'Cause at a minimum, your number crunching becomes real.
38:18 When you look at your data, you know you can trust it, right?
38:21 And context, sure, you can start providing
38:23 the missing context in skills
38:25 by hard coding it for smaller use cases, right?
38:28 That already is a massive efficiency gain.
38:30 Like for example, if you do a weekly meeting
38:32 that looks at all of your revenue across the board
38:34 where you talk about all your top line metrics,
38:36 your new business acquisition, your churn,
38:38 your product usage, your, you know, like expansions
38:41 and conversion rates and what have you,
38:43 that is not that much of a work.
38:45 Like you could, a couple of people could spend like
38:47 two to three weeks of time writing all of that context down
38:50 by just interviewing stakeholders
38:52 and coding it properly in the skill, connecting it and saying
38:55 this belongs to these fields and this is the data.
38:58 Your AI will do a marvelous job of running this
39:01 on a weekly basis for you
39:02 and giving you some really great insights.
39:04 So you start with that.
39:05 But then when the skill piece of it
39:07 starts to get more complex, when you're like,
39:09 oh, I wanna now start adding unstructured data to it.
39:12 I wanna now start analyzing these things.
39:15 At that time, let's say if you still only have 20 calls
39:18 a month, if you're a small business,
39:20 those 20 calls don't need to be vectorized.
39:22 You can feed all of that to the LLM, that's fine.
39:24 But then when that 20 becomes 200,
39:26 okay, let's bring in a vector database now, right?
39:29 And now all of a sudden you've added one more layer to it.
39:31 And then as your entities grow and it becomes more complex
39:34 and now you're realizing that you're holding all of this
39:36 together with a lot of brute force in your semantics,
39:39 at that point, maybe start thinking about a graph database.
39:42 So it's like, you can graduate, but it doesn't,
39:45 so it may sound a little bit like I'm contradicting myself
39:48 where I talked about doing a lot of foundational work,
39:50 but the point here is that you do the foundational work
39:53 before the problem gets really large.
39:55 So don't let it go to 20,000 transcripts
39:58 before you think of a vector database, right?
40:00 But you don't wanna worry about it at 20 either.
40:02 You, when you're still in the middle of understanding
40:06 how useful this is for you, at that point,
40:08 don't invest in tech, just do the work, right?
40:11 Hard-coded in the skill, understand the value it drives.
40:14 Once the value is established, then start doing maybe,
40:18 maybe get like a cloud-based store
40:19 where you pay more per record, that's okay.
40:22 You know, VVA is a great example for a vector database.
40:24 Like if you wanna just start with it,
40:26 not invest too much in it to understand
40:27 whether it's gonna work for you,
40:29 just get a subscription for that.
40:30 Get your vectorized data into it, start using it,
40:34 see how well it goes.
40:36 Now you know you wanna scale it,
40:38 then invest in something that may be on like,
40:40 in your own cloud or what have you, right?
40:42 So it's almost like you're letting your maturity dictate
40:47 what piece of that grand vision you're gonna put next,
40:50 and you're doing it in a way that you start doing it
40:53 before it becomes a problem.
40:54 You have a testing period,
40:56 then you have a early product adoption period
40:59 where you buy something fast but expensive
41:02 on a per record basis but it's still cheap
41:04 on a monthly aggregate level,
41:06 and then you talk about scale.
41:08 That way you've graduated into the curve
41:10 of that tech as you need it, right?
41:12 - That's the way we think about it here too.
41:14 It's making stage fit decisions.
41:17 You don't need to play business at a certain point
41:19 in your company's maturity.
41:21 But like you said, it may be being half a step ahead.
41:24 So hey, you know the scale's coming,
41:26 start working on these things.
41:28 I'd like to dive deeper into vector database
41:31 and just talking about why it's important,
41:35 when you mentioned, hey, at 20 transcripts
41:38 it's not a big deal at 2000,
41:40 it's gonna be really important.
41:41 So talking through those factors
41:44 that start to make it important and why,
41:48 and then technically how to go about implementing
41:52 or anything that teams should be thinking about
41:54 before deploying.
41:56 - Yeah, absolutely.
41:57 It's such a great topic actually, right?
41:59 Because when you have unstructured data
42:04 that contains a goldmine
42:06 of so much sort of qualitative information
42:09 that you may not be getting in your numbers,
42:12 it's almost like a missed opportunity
42:14 to not be able to leverage that.
42:16 LLMs and AI agents make it really easy to do that today.
42:20 But in the early days,
42:22 most of the time what people will do
42:24 is that they will say, here are my transcripts,
42:27 go chat GPT or cloth or whatever,
42:30 find me all of the calls I've had
42:35 where a negative sentiment was expressed.
42:38 So typically what an LLM would do
42:40 is it would just read everything you gave it, right?
42:43 But unbeknownst to you in the backend,
42:46 it's actually sampling
42:47 because it's not gonna go through 2000 records, right?
42:50 So if the record count is small,
42:54 like I said, 20 transcripts,
42:55 it'll read all of it, after reading each one,
42:58 it will understand if it was actually negative
43:01 in any parts of it.
43:02 And if it wasn't negative, it'll discard.
43:04 And if it was negative based on what you asked for,
43:06 it'll keep that.
43:07 Then it will compile the list of transcripts
43:10 that specifically had negative commentary,
43:12 then it will analyze those and then give you the answer.
43:15 Now, imagine having to do that every single time
43:18 and your records aren't 20, but 200, 2000.
43:22 And out of which,
43:23 let's say that the incidence of negative commentary
43:25 ends up 10% of the time.
43:27 Now what you're doing is you're going through 2000 records
43:30 to find only 200 that actually are relevant
43:32 to your inquiry, right?
43:34 That is expensive.
43:36 That kind of compute,
43:37 even if somebody gives it to you for free today,
43:39 the GPU cycles are what they are.
43:42 It takes how long it takes for the system to do this.
43:45 So at a minimum, you're just like,
43:48 you're essentially red lining
43:50 to go to a grocery store, essentially.
43:51 Like if you do the car analogy, right?
43:54 It's not a thing that you want to do
43:57 or is sustainable in the longterm.
43:59 So what a vector database would let you do
44:01 is that it would automatically catalog and index
44:05 all of this conversation data
44:08 into all of the different topics that it could pertain to.
44:12 Which means that when you,
44:14 so think about like a VLOOKUP on steroids,
44:17 but with text data rather than number data.
44:22 So it's like basically if you say negative,
44:26 it will do like a fuzzy VLOOKUP on all 2000 records
44:30 to find the 200 that already have been tokenized
44:33 and sorry, like by been sort of,
44:35 they've been classified in topics
44:37 to say that these had negative.
44:39 And those 200 only are the ones that go to the LLM.
44:43 So the first part of it is not LLM.
44:45 It's very cheap, like traditional software.
44:47 And vector databases have been around forever by the way.
44:50 So, you know, it's just a simple query.
44:54 It runs a query.