1 00:00:00,463 --> 00:00:01,423 My name is Jake Behrens. 2 00:00:01,423 --> 00:00:01,943 I'm with Cirrus. 3 00:00:01,943 --> 00:00:03,383 I'll be helping moderate the room today. 4 00:00:03,383 --> 00:00:04,943 A couple of really brief reminders. 5 00:00:05,023 --> 00:00:10,863 One of them is as we move to questions later, if you could wait for me to come around with a microphone, that'll help make sure that everybody can hear your question. 6 00:00:10,943 --> 00:00:11,903 It is a pretty full room. 7 00:00:12,063 --> 00:00:17,063 So but with that, I have the honor of presenting our presenters this morning. 8 00:00:17,063 --> 00:00:20,943 So I'll go ahead with an introductions here and we can get rolling with their presentation. 9 00:00:20,943 --> 00:00:21,183 So 10 00:00:21,583 --> 00:00:22,463 Good afternoon, everybody. 11 00:00:22,463 --> 00:00:34,223 So it's my pleasure to introduce Zachary Deigert, Senior Data Scientist, Justin Claycomb, Assistant Director of Data Science, and Melissa Hollis, Director of Data Analytics at Principal Financial Group. 12 00:00:34,863 --> 00:00:43,663 So together, they bring a deep experience in designing, deploying, and scaling AI and analytics solutions that deliver measurable business value across the enterprise. 13 00:00:44,143 --> 00:00:56,703 Their collective work spans machine learning, natural language processing, and large language model applications with a strong focus on operationalizing AI and turning complex data into reliable, actionable insights. 14 00:00:57,343 --> 00:01:04,783 So for today, in today's sessions, they're going to share a practical approach to automating data detection, extraction, and standardization. 15 00:01:05,183 --> 00:01:11,103 So demonstrating how AI can streamline data preparation, improve consistency, and unlock efficiency at scale. 16 00:01:11,343 --> 00:01:15,343 So please join me in welcoming Zach Deigert, Justin Claycomb, and Melissa Hollis. 17 00:01:15,343 --> 00:01:24,343 Thank you. 18 00:01:24,343 --> 00:01:26,383 Good morning, or good afternoon, everyone. 19 00:01:26,383 --> 00:01:27,983 Sorry, it's the afternoon already. 20 00:01:28,943 --> 00:01:30,063 So I'm going to get started. 21 00:01:30,303 --> 00:01:31,503 I'm Melissa Hollis. 22 00:01:31,503 --> 00:01:32,303 I, as 23 00:01:32,663 --> 00:01:35,023 I mentioned worked at Principal, I work at Principal. 24 00:01:35,623 --> 00:01:38,063 And I'm going to talk about AI attribute intelligence. 25 00:01:38,103 --> 00:01:45,143 And before I begin, I wanted to read the fine print, and that is that the views expressed here are those of the employees, us, and not necessarily. 26 00:01:45,223 --> 00:01:46,223 Sorry, that was a principle. 27 00:01:48,783 --> 00:01:50,223 We have a short agenda today. 28 00:01:50,503 --> 00:01:52,943 I'm going to start us off setting the stage. 29 00:01:53,023 --> 00:01:57,503 Then Zach is going to talk about the solution and walk through that. 30 00:01:57,503 --> 00:02:01,343 And then Justin's going to finish it up by talking about the value of scaling. 31 00:02:04,943 --> 00:02:12,223 So with regard to setting the stage, I thought of an experience that I had recently, and I thought I would share kind of what that looked like. 32 00:02:12,543 --> 00:02:14,063 But I was planning a road trip. 33 00:02:15,183 --> 00:02:17,663 six days, five nights, three states. 34 00:02:18,063 --> 00:02:21,983 And a part of that is really finding all the restaurants along the way as well. 35 00:02:22,703 --> 00:02:26,863 So I wanted something that was somewhere between Chicago and Two Rivers, Wisconsin. 36 00:02:27,743 --> 00:02:30,943 So I found this really interesting place right off of I-41. 37 00:02:30,943 --> 00:02:33,103 It's called the Mars Cheese Castle. 38 00:02:33,383 --> 00:02:34,703 If anybody's ever been there. 39 00:02:34,703 --> 00:02:35,703 Okay, excellent. 40 00:02:36,023 --> 00:02:36,463 Awesome. 41 00:02:37,583 --> 00:02:39,023 I thought it was a really good choice. 42 00:02:39,023 --> 00:02:40,223 It looked really interesting. 43 00:02:40,543 --> 00:02:43,743 I clicked on the link for the menu and it took me to a PDF. 44 00:02:44,143 --> 00:02:46,863 And this PDF, the first part of it was an order form. 45 00:02:47,663 --> 00:02:58,543 And it suggested download the form, print it out, fill it in, fax it back to them, call with your credit card number to confirm, and then your order could be picked up. 46 00:02:59,303 --> 00:03:03,743 Because I was really looking for a place where I could order ahead of time and then just pick it up as we drove through. 47 00:03:05,063 --> 00:03:10,783 And I thought, well, this probably is not going to be the best option for me because I don't think that I have ever faxed anything in my life. 48 00:03:12,783 --> 00:03:18,783 However, I am very familiar with food delivery apps, right, which are much different and have a much different interface. 49 00:03:20,783 --> 00:03:24,943 So in a food delivery app, a lot of the decision making or the thinking is done for you. 50 00:03:24,943 --> 00:03:28,703 There's a lot of constraints on what you can and cannot put into those apps. 51 00:03:29,103 --> 00:03:38,303 So you often will pick out an entree and there may or may not be check boxes to, you know, add extra sauce or remove the onions or something like that. 52 00:03:38,863 --> 00:03:42,583 Sometimes that decision is not available to you, and so that's what you have to go with. 53 00:03:42,863 --> 00:03:51,823 This is much different from the Mars Cheese Castle, where you could put in, you handwrite, you know exactly what you want, maybe what you want to add, what you want to remove. 54 00:03:52,783 --> 00:03:55,823 But with that comes a lot of interpretation on their end, right? 55 00:03:55,823 --> 00:03:58,463 They're looking at your facts. 56 00:03:58,943 --> 00:04:06,903 that has come through, maybe deciphering your handwriting, figuring out exactly what you want to order, and how much that's going to cost if you've added a lot of things to that. 57 00:04:06,903 --> 00:04:12,143 And then eventually it will get into their system, go into the kitchen, and then get completed. 58 00:04:13,663 --> 00:04:17,663 So 2 very different experiences, both from the front end and the back end of this. 59 00:04:21,183 --> 00:04:30,223 So then I thought about it more and I thought about how a lot of the data that we receive as an organization looks a lot more like that fax than what would come out of the food delivery app. 60 00:04:30,943 --> 00:04:31,223 Right? 61 00:04:31,223 --> 00:04:36,383 We don't have a lot of constraints maybe to what information is coming into our walls. 62 00:04:36,703 --> 00:04:40,943 And a lot of that is because we also want to be really good a business partner. 63 00:04:40,943 --> 00:04:42,223 We want to be good business partners. 64 00:04:42,223 --> 00:04:45,823 We don't have a lot of like requirements on how you interact with us. 65 00:04:46,223 --> 00:04:50,943 We just want you to be able to interact with us in whatever capacity is possible. 66 00:04:52,223 --> 00:04:57,263 So we do, and I'm sure many of you also receive a lot of external information in a lot of different ways. 67 00:04:57,583 --> 00:05:02,863 It can be in a spreadsheet, a Word document, PDF, e-mail, text. 68 00:05:03,103 --> 00:05:05,503 It could be a spreadsheet inside of a Word document. 69 00:05:05,823 --> 00:05:07,503 It can be all sorts of different ways. 70 00:05:08,063 --> 00:05:16,303 But you're really then having to take that additional time to really interpret and translate that information that you've received in order to fit it into the systems that you're using. 71 00:05:21,463 --> 00:05:26,703 And so although you may be receiving information in different ways, there's often generally some patterns in what you do receive. 72 00:05:27,023 --> 00:05:29,423 So I'll use as an example, like date of birth. 73 00:05:29,583 --> 00:05:32,143 This might be a column that you're expecting to be receiving. 74 00:05:32,543 --> 00:05:36,063 And it may be called date of birth, it might be called DOB, it might be called age. 75 00:05:36,743 --> 00:05:43,103 And it might be a month, day, year, or a day, year, month, day, year, either way. 76 00:05:43,103 --> 00:05:44,743 You can get the point. 77 00:05:45,383 --> 00:05:46,703 And it may just be the age. 78 00:05:46,863 --> 00:05:51,583 So there's a lot of different ways that it may come in, but in general, there's some patterns that you're able to look for. 79 00:05:52,543 --> 00:06:01,823 And ultimately, then you have teams or individuals that are reading this information, they're translating it, they're maybe remapping it, and then eventually getting into those systems. 80 00:06:02,303 --> 00:06:06,303 This process is not something that is easy to 81 00:06:07,023 --> 00:06:07,983 pick up, maybe. 82 00:06:08,143 --> 00:06:10,223 There's a lot of nuances in what is happening. 83 00:06:10,903 --> 00:06:15,103 It also is prone to errors, and ultimately it does not scale very well. 84 00:06:19,983 --> 00:06:22,303 So this can look a lot of different ways from the outside. 85 00:06:22,303 --> 00:06:28,703 It's often considered maybe a processing delay, because you have the information, why can't you just process it faster? 86 00:06:29,103 --> 00:06:36,063 But really it's often more of that interpretation or translation delay that's causing the delay. 87 00:06:40,303 --> 00:06:44,783 And this is not a problem just for principal or for our industry, for financial services. 88 00:06:44,783 --> 00:06:46,463 I believe it happens everywhere. 89 00:06:46,543 --> 00:06:56,783 Anywhere you have information that's coming into you from an external partner, it can also even happen internally if you have systems that don't talk to each other well, anything like that. 90 00:06:57,103 --> 00:07:04,143 But any time where there's a lot of data coming in, there's the potential for this translation or interpretation problem. 91 00:07:07,423 --> 00:07:14,063 So then that begs the question of why are humans still acting as translators between messy inputs and rigid systems? 92 00:07:14,383 --> 00:07:18,383 And to answer that question, I'm going to turn it over to Zach, who's going to talk about our solution. 93 00:07:18,383 --> 00:07:18,543 Awesome. 94 00:07:19,983 --> 00:07:22,023 Thank you, Melissa. 95 00:07:22,023 --> 00:07:22,583 Thank you, Melissa. 96 00:07:22,583 --> 00:07:24,783 Hope we're hearing my voice. 97 00:07:27,103 --> 00:07:28,383 And the clicker is working. 98 00:07:28,383 --> 00:07:28,703 Sweet. 99 00:07:29,263 --> 00:07:35,343 Melissa gave a great example, a real-life situation of when messy inputs lead to downstream delays. 100 00:07:36,063 --> 00:07:51,823 The solution we're going to speak to today is kind of a framework, a tiered engine that's able to take a variety of inputs, really doesn't matter what type of file type that we iterate on and improve over time through various use cases. 101 00:07:52,463 --> 00:07:57,423 to try and tackle this in a way that is augmenting that interpretation layer. 102 00:07:57,783 --> 00:08:01,383 We're not trying to completely remove that judgment from the human per se. 103 00:08:01,503 --> 00:08:07,583 We still need to make sure that the inputs are correct before they make any changes or affect a customer. 104 00:08:08,143 --> 00:08:11,903 But we can really augment that with a stack of technology. 105 00:08:14,303 --> 00:08:20,783 When I use the word chaos here, it may be a better synonym is compounding variability. 106 00:08:22,223 --> 00:08:30,783 For my stats nerds out there, each one of these different parameters just adds degrees of freedom for the system that make it more complex. 107 00:08:31,103 --> 00:08:37,023 It compounds with each different type of file that you're trying to use, different structures, different formats. 108 00:08:37,583 --> 00:08:47,343 All those things are just making an exponential problem because every time there's a variability in any one of these, it creates that dimension of difference across all of them. 109 00:08:48,063 --> 00:08:50,463 So we kind of have to take these one step at a time. 110 00:08:51,263 --> 00:09:09,023 and kind of attack them in a way where not only do they impact each other, but they're all kind of handled separately so that depending on what type of file type it is, what type of data attribute we're working with, we understand it throughout the entire flow and so that we can get the correct values. 111 00:09:09,503 --> 00:09:13,503 Overloaded values in context is another tricky piece of the chaos. 112 00:09:13,503 --> 00:09:15,103 Overloaded values is when 113 00:09:15,823 --> 00:09:24,223 Maybe a specific data attribute, I think age and date of birth is a good example, where you can have multiple pieces of information within the same data attribute. 114 00:09:24,223 --> 00:09:27,343 So somebody may give you their date of birth, you require age. 115 00:09:27,983 --> 00:09:36,943 It could be a sentence in an unstructured document that contains multiple pieces of information where context comes into play on how that should be interpreted. 116 00:09:37,263 --> 00:09:41,663 And we kind of have to juggle all of these things before we even get to the data itself. 117 00:09:41,983 --> 00:09:49,663 So we kind of split apart the problem and tackle it with specific tools for each of them before we even do any of the extraction. 118 00:09:52,143 --> 00:09:56,383 So this is where we start to talk about the intake of documents. 119 00:09:57,743 --> 00:10:01,343 This is unglamorous and boring work, honestly. 120 00:10:01,343 --> 00:10:20,943 I mean, whether it's an Excel file, PDF, JSON document, text file, image, God forbid, you really need to make sure you understand what type of file that is before you start anything downstream, because that'll affect what type of modules or processes you apply to that specific document to do the extraction later on. 121 00:10:21,583 --> 00:10:26,863 I mean, when it comes to PDFs, is there certain types of scraping or pre-processing we have to do? 122 00:10:27,423 --> 00:10:34,623 If it's semi-structured documents, even Excel documents, are we having to handle multiple sheets within an Excel document? 123 00:10:34,623 --> 00:10:44,063 Are we having to find that a table is not in the top left corner of your spreadsheet and for some reason a customer has given it to us and it's in the middle of the document for some reason? 124 00:10:44,863 --> 00:10:47,823 Is it an embedded spreadsheet within a Word document? 125 00:10:48,063 --> 00:10:48,863 These are some of these 126 00:10:49,583 --> 00:10:57,263 structural manipulations that need to take place before you even start to look at the attributes and start to look for extraction patterns. 127 00:10:58,063 --> 00:10:59,583 And all of them are kind of separate. 128 00:10:59,583 --> 00:11:02,303 So like I said, this is very iterative. 129 00:11:02,623 --> 00:11:09,823 When we take on use cases with this framework, you know, you may have an understanding of the universe and what types of files you expect to receive. 130 00:11:10,463 --> 00:11:11,743 That could change over time. 131 00:11:11,823 --> 00:11:18,383 And so you want to make sure that you're building upon this framework with new use cases so that as you take on additional ones, 132 00:11:18,623 --> 00:11:27,023 You've kind of built up the knowledge base to handle a wide variety, check that file type, handle it appropriately, and kind of save that off for later use. 133 00:11:27,983 --> 00:11:32,063 But then we move on to the extraction page, which is where I'll be spending the most of my time. 134 00:11:34,623 --> 00:11:35,983 We created this tiered engine. 135 00:11:36,543 --> 00:11:39,103 I think of it as a cake, think of it as anything that's tiered. 136 00:11:39,423 --> 00:11:40,303 Doesn't really matter. 137 00:11:40,943 --> 00:11:58,303 But the idea here is to break apart, like I've said multiple times, individual components that can be handled by different tools and bring the right tool to the problem instead of a shiny, hammery LLM AI model to do all of these things, really fine-tune. 138 00:11:58,703 --> 00:12:06,623 the tech stack so that it's handling the right problem with the most efficient, cost-effective, risk-mitigating set of technology possible. 139 00:12:07,023 --> 00:12:27,743 What this looks like is kind of taking a lot of the behind-the-scenes things that you see when you use documents in your favorite chatbot or copilot and kind of breaking them apart and doing them sequentially and using the less sophisticated, simpler technology when you can and only relying on some of the more sophisticated, high-cost, risky stuff 140 00:12:28,103 --> 00:12:41,263 When it can't be solved by some of the more simple stuff, and so documented text operations, think of PDFs, maybe gets overlooked quite frequently, but there's a lot of formatting or metadata within a PDF document. 141 00:12:41,503 --> 00:12:45,103 Think of things like bolded text, color, location on... 142 00:12:45,223 --> 00:13:13,823 on a document, these are all things that could indicate what the data is. it useful? Is it the type of data input I'm looking for? And getting, using this text operations is to kind of get that structure right, so that you at least are in a good place to go ahead and do some of the actual extraction, which is the following ones. So basic natural language processing, this would be like pattern searches. You know, if you're looking for a phone number, at least here in United States, it's maybe 1 digit, but then we can expect 334. 143 00:13:14,383 --> 00:13:44,063 with some maybe hyphens in between. Some of these RIG-X patterns, really straightforward, programmatic, deterministic, right? If I see that pattern, I'm pretty darn confident that it's a phone number. And so you're kind of taking care of some things that you can with really straightforward processing. You can also, like I said, lean on some of the metadata you got from the document operations. If you understand a given form, it's set up in a way where you have the definition of the attribute, so date of birth, 144 00:13:44,343 --> 00:14:10,223 on the left-hand side, bolded, and then maybe the actual data point itself on the right-hand side, or if it's a spreadsheet with your columns and rows, you can use some of the basic document and text operations to inform the natural language processing. If we can't quite get, and maybe I'll breeze over real quickly, but you can also do some dictionary searching and mapping for keywords and phrases. I'll gloss over that because the next slide dives in a little bit deeper. 145 00:14:10,943 --> 00:14:35,983 But if you can't get through it, get to it with deterministic or straightforward programming or code, you can take maybe one layer of sophistication down into machine learning. Here, maybe your attributes can come in slightly different than expected. Date of birth, DOB, that type of thing. Slight variations of an attribute, first name, 146 00:14:36,663 --> 00:15:00,383 F underscore name. Those slight variations are semantic nuances within maybe the attribute name, but they're close enough that a machine learning model would be able to pick up on, all right, based on my training data, I would expect that this unseen, unknown data attribute aligns to the class of data attribute that I'm trying to detect. This is the next step mainly because it is something that we 147 00:15:00,943 --> 00:15:30,063 can train ourselves, understand since a machine learning model, we have accuracy metrics. It's probabilistic in nature. It's not just outputting tokens. It is making a prediction. And that's something that we can monitor and track over time for performance. It's something that we can also use in a feedback loop. And I'll talk about that in the next slide as well. Named entity recognition is an unnatural language processing concept. Maybe the type of attribute you're going for is a type of thing, an organization, a person, a location. 148 00:15:30,543 --> 00:15:47,103 Those types of attributes are things that machine learning models can predict as well, specifically unstructured text. You can look for that type of part of speech within a sentence, pronouns and whatnot, and you can use that to inform what types of data attributes would align to that. 149 00:15:48,143 --> 00:16:05,983 If we can't do it with those things, yes, we do need to lean on the shiny hammer. Specifically, when you've got sentence structure with some sort of logic to determine the data attribute and you can't grab it directly from the document itself, that's when we leverage a large language model. I'll talk about this a little bit deeper in the next slide as well. 150 00:16:06,383 --> 00:16:28,303 But breaking this out again and not just relying on maybe a chatbot to do this, we have complete control over how we use the large language model and RAG architecture. So what type of chunking or pre-processing should we do to a document before we even ask that question of what data is in this document? We can save money and time, which I'll talk about next. 151 00:16:29,183 --> 00:16:57,423 And then calculations, right? There may be an instance where you've extracted maybe a set of different attributes, but the actual true data point you're going after is a calculation between the different attributes. That's something we'd also handle outside of the use of large language models. Very key here at the bottom, too, once we've run through all this, we've gotten all the attributes we can from a file, we really want to do some data validation steps. For example, age, 152 00:16:57,983 --> 00:17:25,823 can't be negative, right? So you can do some of those really straightforward things to make sure that after I've identified where data attribute is and pulled out the exact values themselves, do they fall within our expected values for that given data attribute? You might be saying, again, great, why not just use the strongest tool for all of these things? And I'll maybe make that a little bit more clear with my double click down into some of these things. 153 00:17:26,863 --> 00:17:53,503 Dictionary intelligence falls into that bucket of basic natural language processing. It's something that is really informed by domain and business knowledge. So this is interaction with stakeholders who have been maybe doing some of that, you know, interpretation themselves of, okay, here's the form that came in, here's the data I need from it, here's how it needs to look for the system. We can take some of those known terms, known data attribute names, 154 00:17:54,223 --> 00:18:21,423 and turn that into a dictionary. And that's a living, breathing dictionary where we use that to curate new examples of potential values within that dictionary to search for and have that feedback loop to say, is this a good one to include, not include? When we include it, is it getting the right attributes out? Is it not? It's trackable, auditable, all that stuff. And really high confidence here because we're just saying, hey, here's the list of different attributes for first name that we've seen. 155 00:18:22,223 --> 00:18:49,423 Here's, is it in this document? If so, go ahead, grab it. If it's not, go ahead and try and use some of the more sophisticated techniques. We use that dictionary to then train our machine learning model, right? So we give it a list of, here's training data, here's what the different variations of first name are, here's a set of variations of first names you've never seen before, and a whole bunch of different data attribute names, try and predict the data attribute names that align to first name. 156 00:18:50,383 --> 00:19:03,503 And get better and better at that. This is again good because it relies on more mathematical probabilities. It's really straightforward. It's converting that text into an embedding. 157 00:19:03,983 --> 00:19:32,143 Then embedding encases some of the semantic information around that text and can carry that over into training. You can also use voting mechanisms here where multiple of these really lightweight, cheap, we're not paying for tokens, we're paying for a small amount of compute to make a prediction. We can use multiple of these lightweight small language models or even just machine learning models on embeddings to vote for which data attribute and then aggregate those predictions together. 158 00:19:32,783 --> 00:20:01,263 to be very confident and reduce the risk of getting misses. It's also a confined scope. When you set up these machine learning models, they're just making a prediction from a set universe, so it can't go off, maybe as some of you have experienced with chatbots, and start just generating completely unrelated text, right? It has to predict within the specific set of data attributes you are looking to find within a given form or document. Lastly, 159 00:20:01,743 --> 00:20:26,223 I kind of hinted on the last slide that we lean on these, but again, maybe another reason why this approach is risk adverse, efficient, effective, as well as makes a lot of sense from a cost standpoint. As we heard the keynote speakers start to talk about cost of using these tools in the future may get kind of crazy. Even when we lean on these tools, we can do so because of the pre-processing and other 160 00:20:27,343 --> 00:20:55,343 tech stack pieces we've used, we can do so in a smarter way, right? Instead of sending that entire document to the LLM and say, hey, here's the data attributes I need. Well, I've done that parsing and chunking to realize maybe instead of this 300-page abstract that you need to now use as input tokens, I just need a paragraph from that, right? Because I've identified, here's the keywords or phrases, a portion of that document that correlates to the data attribute that I'm looking for. 161 00:20:55,983 --> 00:21:24,143 That's where that keyword searching and some of that document processing on the front end kind of comes back in later to interpret. We also have strategies around which model we use. It can be lightweight and cheap even from an LLM perspective. Question answering and RAG solutions don't really require the latest and greatest machine or large language models. And we have parameter selection going on as well. Chunking and different 162 00:21:24,743 --> 00:21:46,183 Brag implementations allow you to kind of manipulate how much overlap and how much different sections you want to look at within a document. Again, we kind of try and avoid that by looking for the specific section with the document that would make the most sense to find that data attribute in. And then your stereotypical prompt engineering and instructions you can provide to the model. 163 00:21:46,623 --> 00:22:12,383 Still here, very crucial that we want to make sure that we're validating any of the data that we pull out at the end of the day before providing it to a human for review or a system for ingestion. Do want to call out these important considerations. I just did one of them at the end of that last slide, but we're making sure that human is involved. Feedback loops on the beginning, making sure that we're using their domain knowledge to guide 164 00:22:13,223 --> 00:22:39,103 our dictionaries, which is really the cornerstone of the rest of the tech stack. And then we want to have validation and quality checks going on for some of those edge cases or examples of ones that do leverage maybe the machine learning or the large language models. We want to provide some level of confidence score accuracy in the past for that data attribute on a source of truth. So they kind of have a feel for how much they should rely on or just take this as 165 00:22:39,503 --> 00:23:09,023 as what it is instead of maybe reviewing and going back to that source document and validating themselves. We do have defined boundaries on all of this work, privacy considerations as well, making sure that we're not using anything outside of the bounds that privacy and risk and governance would be upset with us about. Making sure that anything's redacted or filtered that shouldn't be used and making sure that any data that we have access to is getting removed from people's access when it's unneeded. 166 00:23:10,303 --> 00:23:37,023 And of course, it's a highly governed and regulated industry, so we have a bunch of risk assessment, validation checkpoints, and making sure that we're following those governance standards. I spoke a lot for a long time, but with that, I'm going to pass it to Justin to talk about the value of scaling. Appreciate it. All right. I am going to beg forgiveness. I woke up with the worst sinus infection this morning. So we're all going to walk through this together. 167 00:23:38,183 --> 00:23:45,143 I'm going to do something I promised I wasn't going to do, and I'm going to basically read from my presentation, but just give me a little bit of freedom. 168 00:23:45,823 --> 00:23:50,583 It'll be like the Michael Jordan flu game of reading an AI scaling speech. 169 00:23:50,583 --> 00:23:51,663 We'll see how it goes. 170 00:23:52,223 --> 00:23:57,943 So we've also talked already a little bit about scaling in general. 171 00:23:57,943 --> 00:24:04,463 I mean, we're no Mars cheese castle with one type of ordering process, but Principal has lots of different challenges. 172 00:24:04,703 --> 00:24:08,143 We're a large federated organization with lots of different data sources. 173 00:24:08,743 --> 00:24:11,983 And we're known to be easy to do business with. 174 00:24:13,263 --> 00:24:15,183 for good or for better or for worse. 175 00:24:15,423 --> 00:24:17,503 As Melissa mentioned, we generally accept all incoming data. 176 00:24:18,543 --> 00:24:22,383 And we need to be able to make the data ready for the company to use in the best way possible. 177 00:24:22,543 --> 00:24:23,983 So that's where this process comes in. 178 00:24:24,463 --> 00:24:28,183 That being said, we click on the notice, you know, one, this is a one-off problem. 179 00:24:28,183 --> 00:24:29,583 This happens across the area. 180 00:24:29,583 --> 00:24:32,863 So every one of our BUs has those, business units has those problems. 181 00:24:33,183 --> 00:24:35,823 Different teams, different data sources, different business needs. 182 00:24:36,463 --> 00:24:39,663 The same translation work is showing up over and over and over again. 183 00:24:40,703 --> 00:24:45,983 So that repetition showed us that this isn't just a one-time use case problem. 184 00:24:45,983 --> 00:24:48,783 It's a gap in our capabilities of principle, which we acknowledged. 185 00:24:48,823 --> 00:24:52,383 And it's important to know that we, here, we just want to take pause. 186 00:24:52,383 --> 00:24:54,383 We don't use AI just to use AI of principle. 187 00:24:54,783 --> 00:25:01,903 But when we see that, we say we use it where it's appropriate and it's where it speeds our processes, and again, with governance and human-in-the-loops the whole time. 188 00:25:03,023 --> 00:25:11,103 So instead of trying to solve each request independently using my esteemed colleague Zach Deiger's process, we shifted to solve this problem at a bigger scale. 189 00:25:11,663 --> 00:25:12,703 So how do we do that? 190 00:25:13,543 --> 00:25:26,943 What we really became was moving from individual solutions, doing this every time and doing this process every single time, to really working towards a shared attribute intelligence capability, something that the teams can reuse rather than rebuild. 191 00:25:27,103 --> 00:25:32,463 So what's in the shared attribute intelligence capability in that platform? 192 00:25:33,103 --> 00:25:43,503 I can't give away the entire secret sauce of everything that we do, but it involves shared workflows, standardized codes, reusable attributes, and lightweight AI agents that help bind it all together. 193 00:25:45,343 --> 00:25:48,023 A part of standardizing the workflow is not just the data itself. 194 00:25:48,023 --> 00:25:57,583 We're standardizing how the attributes are defined, extracted, validated, and measured, so new use cases can move forward faster without sacrificing consistency. 195 00:25:59,543 --> 00:26:03,103 What that leads to is the shift, the teams where they spend their time. 196 00:26:03,263 --> 00:26:05,023 It doesn't remove anyone from the loop. 197 00:26:06,143 --> 00:26:10,783 Just less effort goes into the one-off mechanics and wiring of every one of these processes. 198 00:26:11,183 --> 00:26:14,103 More effort goes into data quality, confidence, and reuse. 199 00:26:14,103 --> 00:26:18,143 And those are really the things at principle that go to scale. 200 00:26:18,143 --> 00:26:19,903 Those are the ones that provide value at scale. 201 00:26:20,543 --> 00:26:21,743 So the key idea here is simple. 202 00:26:21,743 --> 00:26:23,063 We're not just solving single problems. 203 00:26:23,063 --> 00:26:28,703 We're building a repeatable pattern that makes the next 10 problems easier, safer, and to handle. 204 00:26:31,543 --> 00:26:38,303 So just to go where we're at, one of the biggest limitations we saw is just how long a day to add anything new to the process. 205 00:26:39,383 --> 00:26:40,983 A small new attribute, a new field. 206 00:26:41,103 --> 00:26:47,503 In the beginning, every new attribute we added to the process turned into a small engineering and ML project, which slowed everything down. 207 00:26:47,503 --> 00:26:51,823 What we're doing here and moving forward as we scale it is we're trying to change that model. 208 00:26:52,143 --> 00:26:55,503 The new attributes are really added through configurations, not custom builds. 209 00:26:56,183 --> 00:27:01,663 That means we're defining what the attribute is and how it should behave without rewriting the engine that we're using every single time. 210 00:27:02,783 --> 00:27:16,703 So, like, with just an example, the old way to add payment frequency, we would write new parsing logic, we'd have new regex or ML code, we'd create new deployment, we'd add new test, new maintenance paths, we'd revisit the code again when there's new variations that appear. 211 00:27:17,183 --> 00:27:19,983 Every attribute became a mini-project. 212 00:27:20,583 --> 00:27:24,463 When we're looking at the new one, we're really talking about configuring the new attribute. 213 00:27:24,463 --> 00:27:32,463 We're looking at, so if we take payment frequency again, we're talking about all the loud values of payment frequency, monthly, quarterly, annual, singles. 214 00:27:33,023 --> 00:27:36,823 And we're talking about also adding in what are synonyms for that? 215 00:27:36,823 --> 00:27:37,783 What are signals for that? 216 00:27:37,783 --> 00:27:40,063 What are signals for that type of variable? 217 00:27:40,063 --> 00:27:42,383 So billed monthly, annual premium. 218 00:27:42,623 --> 00:27:45,983 At principal, there's a myriad of ways to talk about how 219 00:27:46,943 --> 00:27:48,703 premiums and payments and stuff like that. 220 00:27:48,703 --> 00:27:51,903 So we always make sure that we have all the cinnamons available. 221 00:27:52,463 --> 00:27:54,863 The extraction rules, though, we also look at that too. 222 00:27:54,943 --> 00:27:56,503 So he was talking about that here. 223 00:27:56,503 --> 00:27:59,623 And this one was looking at the various places that it could exist. 224 00:27:59,623 --> 00:28:03,103 And that's going to be based on patterns that we know already that exist. 225 00:28:03,103 --> 00:28:07,583 So it's going to go through those various patterns and see which one is the best pattern to look for that particular one. 226 00:28:07,903 --> 00:28:09,263 So it's going to look in structured fields. 227 00:28:09,263 --> 00:28:12,943 If it's present there first, then it's going to go to the free text sections and add it. 228 00:28:12,943 --> 00:28:16,783 It's all based on what we know has happened in the past with those various fields. 229 00:28:17,103 --> 00:28:19,263 It's been part of the engineering that goes into it. 230 00:28:19,663 --> 00:28:26,783 And then for each of those ones, we're also going to have the confidence and the validation tests that go into that one, so they're automated and ready to use for each one. 231 00:28:27,503 --> 00:28:36,143 And then we're also going to flag ambiguous cases for review, so we can spend more time looking at those types of cases than we do engineering in the 1st place. 232 00:28:36,623 --> 00:28:40,623 So under that new scaling, we're not making a new pipeline. 233 00:28:40,703 --> 00:28:42,063 There's no new deployment. 234 00:28:42,143 --> 00:28:42,783 There's no 235 00:28:43,263 --> 00:28:46,623 custom code path, unless it's truly something that we've never experienced before. 236 00:28:47,023 --> 00:28:55,583 It's the same engine that handles policy type and now handles payment frequency because we are telling it what to look for and not to how to rebuild itself. 237 00:28:58,263 --> 00:29:01,983 And a key part of that is every attribute follows the same definition pattern. 238 00:29:01,983 --> 00:29:07,743 We're being explicit about what we're detecting, how it's extracted, how confidence is measured. 239 00:29:08,023 --> 00:29:12,783 A consistency there is what allows us really to scale this whole process, not just given the speed. 240 00:29:13,823 --> 00:29:15,903 And I also wanted to make sure we reference this again. 241 00:29:15,903 --> 00:29:20,383 We're very intentional about using rules first or your ML first. 242 00:29:21,023 --> 00:29:23,463 We will make sure that that's built into the whole process. 243 00:29:23,463 --> 00:29:29,743 At the heart, we really want to follow the guide path that Zach laid out. 244 00:29:30,223 --> 00:29:32,863 If something can be handled deterministically, we're going to do that. 245 00:29:33,743 --> 00:29:35,743 AI really is only layered where there's 246 00:29:36,143 --> 00:29:38,783 vulnerabilities or ambiguity that requires it. 247 00:29:39,063 --> 00:29:43,183 And it keeps everything really predictable and controls for risk. 248 00:29:45,263 --> 00:29:46,383 Also means scaling. 249 00:29:46,703 --> 00:29:48,703 When we talk about this one, we want to scale this one. 250 00:29:48,703 --> 00:29:53,343 We want to make sure that this one works both across structured and free text data. 251 00:29:53,423 --> 00:29:55,143 So we have the same process that works for both. 252 00:29:55,143 --> 00:29:56,703 And that's part of our scaling of this one. 253 00:29:57,103 --> 00:30:03,503 So we don't need to have different approaches depending on the input type, same engine, same guardrails, same onboarding motion. 254 00:30:07,743 --> 00:30:12,063 So the end result is that we add new attributes, then it becomes fast, repeatable, low risk. 255 00:30:12,863 --> 00:30:14,623 That's what makes it usable at scale. 256 00:30:15,823 --> 00:30:22,223 One of the things we wanted to make sure when organizations try to move faster with data, control is usually one of the first things that gets compromised. 257 00:30:22,223 --> 00:30:30,783 People skip validation, they introduce special cases, they very frequently hard code exceptions just to get something to go live. 258 00:30:30,783 --> 00:30:32,863 We wanted to make sure that we're not doing that in this instance. 259 00:30:33,223 --> 00:30:36,703 What we build intentionally is designed to avoid that trade-off. 260 00:30:36,703 --> 00:30:40,783 When new data source comes in, it follows the same repeating path, same onboarding path. 261 00:30:41,423 --> 00:30:47,423 Consistency really allows us to introduce speed without introducing chaos in this whole process. 262 00:30:48,143 --> 00:30:51,743 We have a very intentional governance privacy policies that we follow. 263 00:30:51,743 --> 00:30:58,783 We have a very intentional data validation process that we follow, and we make sure that whatever we're doing to scale doesn't change that. 264 00:31:01,343 --> 00:31:06,303 Instead of building new pipelines every time, we're just reusing what already exists. 265 00:31:09,663 --> 00:31:18,063 This means that we have created expansions that require very little new engineering, even as the volumes and complexities of the data gets increased. 266 00:31:19,263 --> 00:31:24,143 And we just make sure that we're putting the control baked into it, not adding it later. 267 00:31:24,543 --> 00:31:28,143 All that stuff, the confidence scores, it's all part of standard flow. 268 00:31:29,623 --> 00:31:38,143 The practical outcome for us is that we can really move a lot faster with confidence, and we adopt new data sources quickly, scale the platform over time, which is what we're doing right now. 269 00:31:38,143 --> 00:31:44,383 We're still in the middle of kind of scaling the whole thing, and we move, we maintain consistency and trust with our stakeholders. 270 00:31:45,343 --> 00:31:49,303 It's really designed to move fast responsibly. 271 00:31:49,303 --> 00:31:52,543 That's where I'm at. 272 00:31:52,543 --> 00:31:55,703 I need to sit down. 273 00:31:59,903 --> 00:32:00,303 Awesome. 274 00:32:00,303 --> 00:32:05,263 We started a little early, but we did want to leave time for any questions deliberately. 275 00:32:05,903 --> 00:32:06,743 So I'll stand up. 276 00:32:07,183 --> 00:32:09,103 But appreciate it. 277 00:32:09,103 --> 00:32:10,583 And yeah, open for questions. 278 00:32:10,583 --> 00:32:17,993 Who's got the first question? 279 00:32:18,273 --> 00:32:20,353 This isn't a question, it's a statement. 280 00:32:20,353 --> 00:32:23,153 You don't just have to experience it. 281 00:32:27,063 --> 00:32:27,863 and I do want to be clear. 282 00:32:27,863 --> 00:32:30,383 I am still planning on going to the Mars Cheese Castle. 283 00:32:30,383 --> 00:32:31,743 It did look really tempting. 284 00:32:32,063 --> 00:32:33,703 And I've been thinking about cheese curds ever since. 285 00:32:34,063 --> 00:32:34,303 So. 286 00:32:41,893 --> 00:32:42,853 Any other questions? 287 00:32:44,853 --> 00:32:45,333 Behind you. 288 00:32:48,493 --> 00:32:54,533 Just to be clear, you process handwritten notes, Google CRs? 289 00:32:56,693 --> 00:32:58,213 Yeah, good question. 290 00:32:58,213 --> 00:32:59,493 I didn't go that deep. 291 00:33:00,223 --> 00:33:00,583 into it. 292 00:33:00,583 --> 00:33:04,623 But that check file type, yeah, if it's an image, we'll do OCR for sure. 293 00:33:04,863 --> 00:33:09,343 We haven't had yet a use case of handwritten notes. 294 00:33:10,303 --> 00:33:16,703 They're out there for sure, especially if you think about our claims area with doctors and stuff like that. 295 00:33:16,703 --> 00:33:19,743 would be a bridge that when we reach it would be more challenging. 296 00:33:19,743 --> 00:33:25,183 But yeah, if it's an image, we will do stereotypical OCR to get the text off of that image. 297 00:33:34,303 --> 00:33:50,863 So maybe I'm doing one that doesn't feel the full picture here, but you guys get in, I'm assuming, a lot of customer submissions and 298 00:33:52,983 --> 00:33:57,743 You're filing all those submissions through this engine to process all that data. 299 00:33:58,143 --> 00:33:58,383 Is that? 300 00:34:00,063 --> 00:34:09,903 Yeah, I think maybe a good way to think about it is there's bespoke siloed areas that are taking different types of customer inputs. 301 00:34:11,423 --> 00:34:17,263 Having a human, maybe they have an Excel macro that does this a little bit or something like that. 302 00:34:17,983 --> 00:34:19,743 Or maybe they are truly 303 00:34:20,383 --> 00:34:29,663 pulling up a PDF, reading through it, pulling up an Excel file next to it, and making a spreadsheet live that then goes downstream for further analysis. 304 00:34:29,983 --> 00:34:43,463 So the idea here is to create a framework that can kind of adapt to each of those different types of use cases and speed up that and kind of augment is the word I think we kind of like to use, that interpretation step. 305 00:34:43,463 --> 00:34:44,383 So it's not just 306 00:34:45,223 --> 00:35:01,143 That manual, it has a level of technology behind it that helps speed that up and kind of moves us ever so slowly as confident gets higher to what you're kind of describing, where all customer inputs for a specific function are getting handled almost automatically, yeah. 307 00:35:18,943 --> 00:35:20,663 How do you handle different types of PDFs? 308 00:35:20,743 --> 00:35:22,463 They come from different sources. 309 00:35:23,183 --> 00:35:32,063 Yeah, there's a lot of good, actually surprising, open source PDF manipulation tools, readers. 310 00:35:32,423 --> 00:35:35,023 And we experimented with a wide variety of them. 311 00:35:35,663 --> 00:35:37,783 We kind of have two that we really love. 312 00:35:37,783 --> 00:35:40,543 I'm sorry, I'm not remembering off the top of my head. 313 00:35:41,983 --> 00:35:45,263 But we Python code does a lot of that work, really. 314 00:35:45,583 --> 00:35:51,903 Open source frameworks for scrubbing, like I said, some of that metadata off of the PDF documents, if I've identified it's a PDF. 315 00:35:52,303 --> 00:35:53,183 But you're exactly right. 316 00:35:53,183 --> 00:35:57,663 Some, like if you open a PDF, sometimes you'll see some, you know, you can click a check box. 317 00:35:57,663 --> 00:35:59,023 That's a different type of PDF. 318 00:35:59,023 --> 00:36:01,503 Or you can highlight the text or you can't highlight the text. 319 00:36:01,783 --> 00:36:07,903 All of those are little intricacies that we try to do the most least sophisticated thing. 320 00:36:08,383 --> 00:36:14,423 If it's a true PDF, we can use those open source packages to scrub pretty much everything off of them. 321 00:36:14,423 --> 00:36:15,823 Thank you. 322 00:36:15,903 --> 00:36:33,353 We don't have any other questions. 323 00:36:33,353 --> 00:36:34,793 Any other closing comments for us? 324 00:36:36,153 --> 00:36:37,593 No, I just appreciate the time. 325 00:36:38,073 --> 00:36:41,513 I will, I think we'll all be at the facilitated discussion potentially. 326 00:36:43,113 --> 00:36:45,993 Come with any questions that maybe didn't come to mind quite yet. 327 00:36:45,993 --> 00:36:49,513 Yeah, I appreciate being here and for all of you giving us your attention. 328 00:36:50,623 --> 00:36:51,023 Thank you.