How AI Is Transforming Unstructured Data from Chaos into Strategic Asset – with Kyle DuPont, Ohalo

Shownotes

Only 29% of organizations really know where their unstructured data lives – while 79% are confident they can extract value from it. Kyle DuPont (Ohalo) and Carsten Bange discuss BARC's latest report on unstructured data: why classification, metadata and AI readiness are the foundation, how to govern data across on-prem, hybrid and multi-cloud, why agentic AI removes the RAG curation burden, and which criteria really matter in tool selection.

Read the report for free: https://barc.com/webinare/unstructured-data/

Kyle DuPont on LinkedIn: https://tinyurl.com/urhzc8ds Carsten Bange on LinkedIn: https://tinyurl.com/37sdzd2s BARC on LinkedIn: https://tinyurl.com/4j96bfnf Stay up to date with our newsletter: https://tinyurl.com/3ft3vpxv

Transkript anzeigen

00:00:00: I almost view AI as a sequel type of query language for unstructured data.

00:00:06: It's not exactly that, but it's like we had all this data that couldn't be queried at all and now we have AI who can query everything and understand them semantically.

00:00:25: Welcome to the Data Culture Podcast!

00:00:27: I'm Carsten Bange, Founder & CEO of BARC and my guest today is Kyle Dupont.

00:00:32: He's founder and CEO of O'Halo, focused and dedicated to unstructured data.

00:00:41: We will discuss the findings of The Latest Barc Survey & Report on Unstructured Data, which really uncovered a lot of interesting insights!

00:00:52: For example we'll talk about the confidence gap meaning that a lot companies are very confident they can derive value from unstructured say they actually know where it is and how to access.

00:01:08: So we talk about, How To Change That?

00:01:11: How To Get Better Access to Unstructured Data?

00:01:14: How to get the governance in order And What Kyle Actually Sees In Companies How They Can Improve The Handling For example also efficiency in handling unstructured data.

00:01:27: We have a special focus on AI so looking at unstructured data, obviously is right now to prepare data for AI or have AI-ready data.

00:01:40: Enjoy the episode!

00:01:42: Hey Kyle great to have you on The Data Culture podcast.

00:01:46: Thanks for having me really pleased to be here.

00:01:48: Kyle we got in touch because you were involved with our Unstructured Data report and that's area of expertise.

00:01:57: and let me ask you first.

00:01:59: What did you found most interesting in terms of results?

00:02:03: Of that report?

00:02:04: Yeah, it was a super super interesting report.

00:02:06: I mean A lot of stuff Was stuff that we knew about but a lot of Stuff is also surprising as well.

00:02:12: so i think That the high level Is that were still any early innings using unstructured data at scale on the enterprise And The report elucidated a lot of where we are now and kind of, where we need to get too in order to become you know into...in order to leverage the intelligence that a general enterprise needs to use.

00:02:35: uh in the AI world.

00:02:36: Um I would say in particularly-the bits that i found super interesting were probably two or three items?

00:02:43: The first was that it was any kind of external forcing function.

00:02:52: It's somebody how that rolls out, right?

00:02:55: If you look at privacy or security regs You see basically policies being created people being hired to manage that and then the process That is kind of defined in those policy gets baked into the enterprise architecture.

00:03:08: The it architecture.

00:03:10: So we definitely saw on the report a The kind of front The front line of AI deployment in the enterprise is really around the policy and kind of though the process deployment right now And then we're kind of lagging behind to be actual Enterprise deployment still.

00:03:28: but I think if we were do the same survey next year, um.

00:03:31: I would be surprised If that's still the case you know?

00:03:32: I think there's a very high Momentum towards going too to kind of consolidating baking end those enterprise requirements into the actual it architecture.

00:03:42: Um so that was number one.

00:03:47: The second one was around just how big the confidence gap is from reality.

00:03:54: So when we asked the question, you know... How confident are your organization can effectively extract value for unstructured data?

00:04:05: There's about a seventy-nine percent confidence level in that.

00:04:08: but do you actually know where your unstructured information is?

00:04:13: You're confident to extract it?

00:04:15: only about twenty nine percent of people uh fully agreed with that statement.

00:04:18: so there's definitely a gap between where people think they are and where they actually are.

00:04:24: That needs to be bridged, right?

00:04:27: And then I kind of you know i think um leads me on what we always knew when were in this world.

00:04:35: but In order to understand your data is You have to be able to classify it Right!

00:04:42: what type of data it is.

00:04:43: You have to understand the data quality principles around that and provide those really fundamental foundational elements about, you can do all the rest of this stuff.

00:04:52: So I think Those are some interesting parts That we either already knew About or are representative Of where were in world today.

00:05:02: Before digging deeper What actually unstructured data?

00:05:05: Could give us a good definition?

00:05:07: Yeah so Unstructured Data Is All The Data you know and care about, deal with every single day right?

00:05:15: It's files mainly.

00:05:17: Its free text in databases even You have the semi-structured world which is lots of JSON objects things like that where you don't a priori already know this schema often.

00:05:32: So yeah unstructured data... The vast majority, eighty percent of enterprises' data are unstructured.

00:05:39: The vast majority is files in SharePoint sites and through buckets, file servers data lakes that type of thing.

00:05:46: Interesting.

00:05:47: so let's look a little bit more into detail on some of the topics you mentioned.

00:05:53: First of all my general question would be why do think companies lagging behind using unstructured data or creating value with it?

00:06:02: What what's your problem here?

00:06:03: Is there like technical problems?

00:06:05: Or maybe business side doesn't see them the value in it?

00:06:09: No, I think...I don't know if they are lagging behind.

00:06:14: What is lagging?

00:06:36: It's not really a very measurable type of productivity.

00:06:40: And at the same time, a lot people are using like RAG systems that actually pretty hard to build to be perfectly frank.

00:06:45: Like RAG is pretty exquisite right?

00:06:48: You have to really carry out data you're using in RAG system and can't deploy it on a billion file scale but maybe a couple hundred files or thousand file scales works well But it's really hard to scale out.

00:07:02: And now, I think as we're entering twenty-twenty six, you know, Agentec AI is being still... The enterprise was like eighteen or twenty four months behind the state of art in Silicon Valley but You started seeing these kind of green shoots of agentec ai coming up and i think that will be... We can talk Garshton about what agentecai actually means phase and really delivering a lot of productivity, a lot value to enterprises in that very kind of scalable way.

00:07:34: That was probably more scalable than what you can actually achieve with RAG.

00:07:38: Given these vast amounts of unstructured data how can I process or prepare unstructured for AI?

00:07:47: Yeah!

00:07:48: That's super important.

00:07:50: so we see customers who have gone To really painful processes, right?

00:07:56: I mean anybody that works in an enterprise knows That you end up like i want to use data for blah-blah Use case or application.

00:08:04: Right it takes you.

00:08:05: You know You have to go through compliance you have to get your security.

00:08:10: Um, right and if you're doing all this manually, you know It's going to take you a month to find the data you have.

00:08:15: Go to the data custodians.

00:08:16: data owners actually understand what data is there.

00:08:19: In The first place you Have to curate that.

00:08:21: um after you have to then go into your security compliance, people will say hey this is the data I actually want use.

00:08:25: and they're going ask well.

00:08:28: Is it applicable?

00:08:29: What are you gonna do to mitigate these risks?

00:08:31: all these types of questions just like enterprise process that goes on that protects the organization right?

00:08:36: um so if you can shortcut that's a big win right.

00:08:40: You got from what could be six few months even years sometimes down till days where we can say Actually We Have an Index Of Where All My Files Are Where All my Instructure Data how it is responsive to security, how it's responsive to compliance policies.

00:08:55: How its relevant or timely to this particular use case that I'm using and all of the metadata serves to feed into a huge speed up deployment in your use cases.

00:09:06: Yeah!

00:09:07: And you mentioned this confidence gap.

00:09:10: so they think can extract value.

00:09:12: but Then you see the governance topics seem to be unsolved, especially starting with a very basic thing finding it or knowing where is.

00:09:21: How does that come?

00:09:22: Do we have any explanation for this big gap?

00:09:25: everybody's had the experience of like the magical I upload a presentation or I upload an Excel file to quad or to chat GBT and they're like wow!

00:09:34: You know i did something with AI And then... That's something that is great, but in fact In order to scale it out you need to like build governance around It.

00:09:47: Like You can't just have everybody in the enterprise.

00:09:50: if for instance we were I was Just in Japan last week and We're talking about a big manufacturing company.

00:09:58: there has around three hundred thousand people worldwide And they are trying To deploy.

00:10:03: on their case i believe it Was Claude deploy Clawed and an enterprise like that, make sure everybody is using it effectively.

00:10:13: Like you're not bringing to me tokens You are actually governing what data can go into AI.

00:10:19: It's a really hard problem to solve.

00:10:21: In essence in their case they want to make sure manufacturing processes are optimized at the various facilities They have.

00:10:30: But if we just let everyone go willy nilly I think You know, let me... I'm getting off the pace here.

00:10:39: Let's re-answer your question if you don't mind Kirsten Yes so why is there a big gap between where people think they are and actually today?

00:10:56: Really it just comes down to the fact that people have experienced this magical moment of uploading an Excel file or presentation to open AIAH or chat to your cloud and had done something with.

00:11:07: But then when you actually get to like enterprise scale applications that require ingesting large amounts of unstructured data, the processes just aren't there around it.

00:11:16: And it starts again as I said with classification and understanding AI readiness of data –the quality.

00:11:31: We also had a question, the study around where this unstructured data is and we could show that it's really... It lives across on-prem hybrid multi cloud environments.

00:11:40: Does that also play big role in terms of governance?

00:11:44: Yeah huge, huge role.

00:11:46: you know.

00:11:48: I was actually at a client office in Japan last week.

00:11:55: I think it's actually not that different than any other IT system in a large enterprise and its just kind of the refrain we've had from, I don't know ever since computers were being used by enterprises.

00:12:06: Enterprises on the IT side are just a cobbled together collection of companies people have bought or servers that people decided to deploy or services they decide to deploy at any kind of large enterprise scale you're going need deal with.

00:12:24: you know cross across cloud, you know multi-cloud hybrid systems right on prem etc.

00:12:31: and And that's just how it is.

00:12:34: And one of the things that this client was saying was like You know How do we actually make sure that?

00:12:37: We have a common governance structure across everything.

00:12:41: So I can ensure That you know i'm being i'm delivering my own structured data in a secure way In a well governed Way and II told them basically You Can try to eat that elephant if you want Right now.

00:12:52: but you Know your your president is saying, hey we need to deliver this by October.

00:12:58: And the main thing I want to say everybody in this world it's just like start small right?

00:13:04: Like if you can build an ROI use case that delivers real value for a single part of your organization... That's where usually really starts trying to develop a governance model thats gonna work across and their cases they're at the extremists but across.

00:13:18: three hundred thousand people is going to take a very, very long time.

00:13:22: But in order to buy-in for that and get budgeted you want really start with the small stuff And I think it's not even about AI.

00:13:31: Data people are always trying kind of provide wide ranging governance structure.

00:13:37: That's an noble task but we need to deliver value within individual teams because when they when you have people singing your praises to other people, that's when you'd have the snowball effect of delivering real value for the rest of the organization.

00:13:56: I'm starting trying to get my head around this very low visibility into unstructured data that many companies seem to have... What's driving it?

00:14:06: Is it fragmented repositories lack of metadata ownership...?

00:14:10: It is also just like nature of unstructured and the world where they had databases, well in structured data world you have databases.

00:14:22: And even if you have thousands of databases and there's hundreds of schemas in those databases it's still a pretty grokable type amount of metadata that you're looking at but... In the unstructured data world there's no such thing as a schema, right?

00:14:41: There is basically metadata around the file and then there are the semantic context of the file itself.

00:14:45: And maybe some tokens that you want to pull out at that via more traditional approaches but... You're looking at like billion files scale across for instance twenty thousand SharePoint sites or bunch of file servers S-three buckets, Databricks whatever happens wherever your own structured data happens to be stored.

00:15:03: Each of those billion files could have in our case in Data Actuaries we pulled back probably on average, around fifty different metadata points for each different file.

00:15:14: So you could have like just a fifty billion data points that your looking at.

00:15:19: and actually the thing an agent needs or person needs in the unstructured data world is actually very small subset of it.

00:15:25: so... You have to go from a world where we don't know anything about the unstructured data into a world when there's an index right?

00:15:32: I mean its very much like the encyclopedia.

00:15:35: if you're looking for the article about mermits, you don't start in a section of the encyclopedia and just kind of scroll through the whole encyclopaedia.

00:15:46: You go to the M-section then find the mermit article that your looking at right?

00:15:51: That's what we need do on the unstructured data world.

00:15:52: so it is also going.

00:16:03: So I think we describe the problem.

00:16:05: A lack of discoverability, visibility and what would be your strongest reasons why companies should invest in especially their unstructured data world?

00:16:17: Ultimately i think that's where company intelligence is... That's the thing they can really leverage to be competitive on market because you know is available to an individual and organization.

00:16:34: It's like the intelligence in know-how, decision making processes that a company has... That why enterprise stays competitive.

00:16:43: And you think about where it was defined?

00:16:45: And also largely unstructured data.

00:16:47: To give you example one of our clients as contractor they do work for government entities and they'll do proposals, right?

00:16:57: And when you do a proposal well how would you get an agent to do a propose.

00:17:02: Let's say you could have an agent build a proposal for every single line item in our government budget Right?

00:17:08: Well uh You have to know a lot of things To Do.

00:17:10: That Like if I was human building that i Would want to Know Things About like Past Proposals.

00:17:14: We've Done um Success Stories we had.

00:17:17: I Have to convene with my finance team to understand whether the proposal is gonna actually make us a profit or not.

00:17:23: Um information about RFPs that are out from the government.

00:17:27: So there's like a lot of different just PDFs, PowerPoints Excel documents that go into that essentially human does now and does it very painstakingly right?

00:17:38: It takes a long time to build that kind of thing.

00:17:39: so with data actually they were able combine that using an agent and build into something, you know.

00:17:57: That's an output right?

00:17:58: But they can deliver to their client.

00:17:59: Yeah makes a lot of sense.

00:18:01: so You do see or you can't see the value in it.

00:18:06: now What what Do companies need to address?

00:18:10: yeah So we already talked about a few issues.

00:18:12: but maybe if you want to give an overview that company say okay I understand the problem.

00:18:17: I

00:18:18: think that the first thing you would want to address is defining what your use case.

00:18:26: How can we deliver something right now, in a next three months?

00:18:31: Build an approvable use-case with value for somebody.

00:18:35: there's just like vision of what you wanna do.

00:18:40: but then after what do I care about in that case?

00:18:46: And then you get into things like AI readiness, right.

00:18:51: Accuracy consistency business relevance completeness timeliness all these whether or not it's labeled properly from a security perspective All of the other questions so that you can say okay now The AI can understand this metadata That wasn't there before Right Like if you don't have any metadata.

00:19:09: If go to a SharePoint site for instance You're Not going To be able to query SharePoint for like how accurate is this data?

00:19:15: You know, you can't ask Like How relevant Is This Data?

00:19:17: It Just Knows About.

00:19:18: If you search in SharePoint it just knows about keywords or maybe like titles of files and things like that file names.

00:19:24: So you have to build this metadata up And then you Have To Say Okay Well In That Case I Need Actually Do Deep Classification Of The File To Understand Not Only That Metadata Rapper Around the File But Also The File Semantic Context.

00:19:36: Together Translates Into Those AI Readiness Points That I Was Talking About Like Accuracy Consistency.

00:19:42: if you can deliver that in an index, then AI could read through MCP or something else like that.

00:19:47: Then you end up with really powerful to deliver those use cases against the vision defined at the beginning.

00:19:56: Do you see patterns of groups and use cases especially benefit from more or better unstructured data?

00:20:05: I think in the unstructured dataworlds security related use cases, and that's kind of what unstructured data was before is like how do we actually just protect this data?

00:20:16: And or How Do We Comply With This Data.

00:20:19: Um...and we've seen you know one of our clients has a big bank!

00:20:22: They started using Data X-Ray for the purposes of compliance and security

00:20:28: right?!

00:20:28: How do we understand whats in files so that we can report to our regulators?

00:20:33: uh..how are managing our records?

00:20:34: essentially then In their case?

00:20:36: but y'know That metadata is there once we've done it, and you can actually use that for AI.

00:20:41: So then they can actually translate that into a variety of AI-use cases that their working on in particular around... They do a lot at corporate banking but KYC refreshing to understand if there's any kind information in this document the KYC records for this particular entity that we're working with.

00:21:08: And it's things like, you know, ten K documents and things like that they'll store.

00:21:12: but there were also scanning for other security purposes.

00:21:16: so... This metadata your generating is kind of what I'm saying useful a lot different use cases in the organization.

00:21:28: But you said that security and compliance are big drivers.

00:21:32: Yeah It's where, if you think about like what the original send of unstructured data was.

00:21:39: The personal computer right?

00:21:40: Where You gave people the ability to just write random data into a computer And that persisted from I don't know Like the seventies until dearly two thousands and then People started getting hacked in.

00:21:52: security of files became A really big problem.

00:21:55: so you had the security team That Was basically the owners Of unstructured Data but now an AI we have finally a way to interface with unstructured data and understand semantically what's in there, right?

00:22:06: And so that's why you typically see the arc of unstructured-data going from security into data as we are here in twenty six.

00:22:18: But like... That's fast changing.

00:22:20: I think There is convergence between security teams.

00:22:24: AI Data Teams Compliance Teams all need really work together In a strong way to leverage all the metadata that each is generating into use cases, because if you're gonna spend a lot of money to scan through a bunch of billion files.

00:22:40: You should use that metadata as much as you can for whatever use-cases you can.

00:22:44: so I hope gives ya like little bit context on why security it's more historic reasons rather than... It just has an historical reason where we are today.

00:22:58: Now you mentioned multiple times AI being like the driver of all this.

00:23:03: Yeah, these are new use cases companies want to implement and since we're talking about language with Sam obviously then unstructured data comes into play here.

00:23:13: on The other hand what?

00:23:14: You just described in terms of things you need to do.

00:23:17: I'm pretty sure AI and LLMs also support that Like they create off metadata for example could give us an overview of that?

00:23:25: yeah It's.

00:23:26: it's amazing change to products like ours and to customers in general, right?

00:23:36: So let's say we were sitting before ChatGPT came out.

00:23:46: We're sitting there and want know what is in a document And the only real technique available was regular expressions keyword searches you know, named it C recognition that type of more traditional machine learning.

00:24:02: Or or You could like build a very custom model for particular particular documents but It was all really hard to do and you just end up with a bunch of false positives Like you didn't really Know too much beyond okay.

00:24:16: there might be credit card number There may be person's name in this document And so naturally limited the types of use cases That you can address.

00:24:24: But now I almost view AI as having almost like a sequel type of query language for unstructured data.

00:24:35: It's not exactly that, but it's like we had all this data that couldn't query at all and now we have AI that can query all of it and understand it semantically to create structure out of unstructured Converging unstructured data and structured data into the same data pipeline, right?

00:24:58: That is And we do that with our clients all the time.

00:25:01: but integrations with various you know Databricks or whatever.

00:25:04: You know data as soon as like wherever data analytics platform are using.

00:25:10: But it's a real big.

00:25:11: It's a huge change from where we were I mean We started our company what eight years ago before this all happened.

00:25:17: And it's A Huge Change From Where We Were Then To Where We Are Now to be able to deliver this type of power for our clients.

00:25:26: You have what's in essence, just a bunch of different thousands of interns that you can deploy at any time.

00:25:32: you want... ...to look all your files now and actually semantically understand them whereas before you didn't have that right?

00:25:39: Like you couldn't scale that so by nature it was limited on the use cases or the scale of these case you could address.

00:25:48: That makes a lot sense.

00:25:49: So very interesting.

00:25:51: Since we're talking about AI, obviously you need to talk about agents.

00:25:54: so that's the new thing companies are adopting.

00:25:57: how do you see their influence on unstructured data and twice versa?

00:26:02: You know if let say you wanted a query bunch of documents To have an outcome right like that proposal use case I was talking before.

00:26:08: Like i want it to Have A rag system or some sort Of system That can query a bunch of Documents.

00:26:19: Well, if I was going to do that you know a year or two ago i would probably create some sort of like analytics pipeline.

00:26:27: Or rag system that ingested a bunch of former proposals maybe a bunch Of like former excel documents whatever happened To be and You Would have to first curate That data.

00:26:41: so i Have A billion files.

00:26:42: i need to Curate The Data in those files to understand whether it's relevant and then I would want to get it down till like the top hundred files or you know, a thousand files that are relevant.

00:26:54: because yeah.

00:26:55: In a rag system You can't put like a billion files into our rag system Because oftentimes It'll just return like the Top K kind of the top fifty results.

00:27:03: And if your end-of-billion files its statistically less likely That you're going to return The right response.

00:27:09: so the right chunk versus Whether Versus like a several hundred or several thousand chunks of files that you're putting into a RAG system.

00:27:18: So, You ended up with these very highly curated systems and in the agent world... ...you don't actually need to know this document is generally relevant for this query And we are allowed to ingest it.

00:27:29: Then the agent will literally just say okay I'm gonna- This type of word within a PDF and it could be on page seventy-two.

00:27:40: And literally, I'll just do a grep of that keyword onto page seventy two.

00:27:47: basically create custom Python code in the fly to parse document find right table or whatever you're looking for there.

00:27:55: build like little custom throw away.

00:27:57: they don't even applications are scripts.

00:27:59: agents look through files for you until agent actually doesn't need all that harness, they just need to actually know what data they can query for and then actually query before that data pull the file back.

00:28:13: And away it goes.

00:28:15: so basically just simplifies everything too much right?

00:28:21: It's a complete shift in how people could build out AI enabled applications or workflows within enterprises.

00:28:30: Let's maybe try to wrap it up a bit.

00:28:36: If you want to give companies the recommendation, so how can they derive most value from unstructured data?

00:28:43: What would be your key recommendations here?

00:28:47: like I said start small And this is a key recommendation for data.

00:28:51: people No matter no matter if its AI that they're trying do or anything else.

00:28:55: starts mall build build a great ROI with a case study internally that you can then go shout from rafters to everyone.

00:29:06: So, that's number one.

00:29:08: But in the pure unstructured data world I think it's clear from the study right?

00:29:13: The things that you need to do are To classify the data first understand what you have in the first place and start building those AI readiness metrics around it.

00:29:22: You know Validate if its good is at the right format.

00:29:25: You need to catalog get unique happen index of it And that's just, it's very clear from the study.

00:29:30: So I would recommend reading this study around that.

00:29:34: as far as tool evaluation The thing that was interesting for me wasn't actually the kind of top things there which is reliability security and performance.

00:29:48: Yeah obviously you want within a tool?

00:29:59: for the agentic world, right?

00:30:02: Because really what we're talking about is like can you actually share data and get data out of the tools that your deploying.

00:30:10: Right And Is That a Good Developer Experience?

00:30:12: Can Agents Use It?

00:30:16: You are going to hear A lot.

00:30:17: if you haven't already heard The word Headless Applications.

00:30:21: it's basically just marketing spend in the agentics world For APIs are developer friendly that you can actually develop against.

00:30:30: And I think, another thing when you're making those decisions to deploy... You really need to focus on.

00:30:38: is this application open?

00:30:40: Can you pull the metadata back?

00:30:41: Is it valuable for...?

00:30:45: Can we get value out of app and provide other tool or agent itself trying build a use case around it?

00:30:52: Excellent!

00:30:53: Thanks so much Kai.

00:30:55: Also an interesting discussion around unstructured data, how we can make use of it and what's happening with all the developments in AI.

00:31:03: So I think that was super insightful.

00:31:05: Thanks so much!

00:31:06: And i wish you all the best.

00:31:07: Say bye-bye.

00:31:08: Hey cheers Kerstin thanks.

Neuer Kommentar

Dein Name oder Pseudonym (wird öffentlich angezeigt)
Mindestens 10 Zeichen
Durch das Abschicken des Formulars stimmst du zu, dass der Wert unter "Name oder Pseudonym" gespeichert wird und öffentlich angezeigt werden kann. Wir speichern keine IP-Adressen oder andere personenbezogene Daten. Die Nutzung deines echten Namens ist freiwillig.