1 00:00:00,009 --> 00:00:01,369 thank you for being here this afternoon. 2 00:00:01,369 --> 00:00:02,809 So my name is Jake Behrens. 3 00:00:02,809 --> 00:00:07,209 I'm a member of Cirrus, so I have the honor of moderating this room this afternoon. 4 00:00:07,209 --> 00:00:07,529 So 5 00:00:07,849 --> 00:00:11,849 I'm going to go ahead and do our introductions here, and we'll go ahead and get rolling with this presentation. 6 00:00:11,849 --> 00:00:13,129 So good afternoon. 7 00:00:13,129 --> 00:00:19,289 It's my pleasure to introduce Matt Vincent, founder of Source Allies, and Ben McCone, staff engineering consultant at Source Allies. 8 00:00:19,289 --> 00:00:29,289 So Matt leads the consultancy focused on data and AI with multiple generative AI systems in production, delivering measurable business results. 9 00:00:29,289 --> 00:00:37,729 So Ben specializes in deploying advanced AI systems with a strong emphasis on reliability, evaluation, and building trust in real-world applications. 10 00:00:38,009 --> 00:00:49,049 So together they work at the forefront of helping organizations move generative AI from experimentation into scalable production-ready systems. 11 00:00:49,049 --> 00:00:58,929 So in today's session, they're going to explore how to close the learning gap in generative AI, demonstrating how systems can continuously improve through feedback, observability, and 12 00:00:59,369 --> 00:01:02,089 optimization without relying on fine-tuning. 13 00:01:02,169 --> 00:01:05,289 So please join me in welcoming Matt Vincent and Ben McCone. 14 00:01:07,649 --> 00:01:07,889 Thank you. 15 00:01:07,889 --> 00:01:08,409 Appreciate it. 16 00:01:09,289 --> 00:01:18,809 So what we're talking about today is something that Ben and I were really excited about when we first saw this paper come out from Stanford. 17 00:01:18,969 --> 00:01:29,449 And then actually while we were developing the talk, a super simple library came out that allows you to do what we're talking about. 18 00:01:29,929 --> 00:01:31,209 It's really exciting stuff. 19 00:01:31,769 --> 00:01:32,649 We're happy that you're here. 20 00:01:32,649 --> 00:01:34,969 We're happy that you weren't scared away by fine-tuning. 21 00:01:35,929 --> 00:01:40,569 because it is kind of a technical topic, and what we're talking about is more approachable than that. 22 00:01:42,649 --> 00:01:52,409 Source Allies, where Ben and I are from, is an IT consultancy based in Des Moines, and we specialize in data and AI. 23 00:01:52,729 --> 00:01:59,049 We've been around for almost 25 years now, and we are proud to be builders who teach. 24 00:01:59,049 --> 00:02:01,449 We do a lot of building, a lot of building of Gen. 25 00:02:01,449 --> 00:02:04,329 AI systems, and we've run into this roadblock that we're 26 00:02:04,809 --> 00:02:06,889 going to talk about today and how to get past it. 27 00:02:08,009 --> 00:02:13,849 And then as we're working, we really love teaching and learning, leveling up together, whoever we're working with. 28 00:02:17,369 --> 00:02:27,529 Back in 2002, I founded Source Allies, and I was one of several teammates who started our data and AI practice almost eight years ago now. 29 00:02:28,249 --> 00:02:30,769 And way back when 30 00:02:30,769 --> 00:02:40,729 When I was in college, my dad led a neural networks research group, and I was what the keynote speaker called today one of the doubters. 31 00:02:40,809 --> 00:02:47,849 I did not believe that anything he was doing was actually possible, but here I am today, Ben. 32 00:02:48,169 --> 00:02:50,009 Yeah, my name is Ben McCone. 33 00:02:50,169 --> 00:02:52,489 I am an open source contributor, a staff. 34 00:02:53,209 --> 00:02:54,569 I'm a consultant with Source Allies. 35 00:02:54,889 --> 00:02:58,329 I've actually contributed to a lot of the tools that we'll show you today. 36 00:02:59,049 --> 00:03:09,049 I have taken multiple clients from idea concept to production with scalable single and multi-agent systems over the years, all of the buzzwords. 37 00:03:09,609 --> 00:03:13,649 I've seen that progression over the last three or four years that we've been doing AI. 38 00:03:13,649 --> 00:03:17,129 I'm really excited to be giving this talk to everyone today. 39 00:03:17,729 --> 00:03:18,249 Back to you, Matt. 40 00:03:18,569 --> 00:03:19,049 Thanks, Ben. 41 00:03:19,609 --> 00:03:35,849 I was actually in San Francisco with Ben at a conference where he was speaking at, and the kind of leaders of these open source AI framework groups, like LangChain, if you do AI development, you're familiar with these, LangChain, DSPy. 42 00:03:36,729 --> 00:03:38,889 Everybody was super excited to see Ben. 43 00:03:38,889 --> 00:03:43,929 It was really weird, like we show up from Iowa and it's like, Ben, thank you for helping us. 44 00:03:44,329 --> 00:03:46,649 So there's some credibility up here. 45 00:03:48,329 --> 00:04:12,009 So what we're talking about today is a report that came out from MIT late last year talking about the state of AI in business, and it talked about this big problem that everybody's running into, this big roadblock that is the AI applications, the generator application that we're building, aren't learning. 46 00:04:12,809 --> 00:04:14,729 You interact with it, you 47 00:04:15,529 --> 00:04:16,889 Tell it, no, that's a little bit wrong. 48 00:04:16,889 --> 00:04:18,649 This is how we work. 49 00:04:18,649 --> 00:04:23,049 Or you want to guide it a little bit more before you expose it to a larger audience. 50 00:04:23,689 --> 00:04:29,369 And it's just not learning from how you actually work within your organization. 51 00:04:30,089 --> 00:04:42,889 Maybe it's because of the things that you want to capture really aren't written down anywhere, but they come out through the interaction of a large group working together with the AI. 52 00:04:43,609 --> 00:04:44,249 This 53 00:04:44,689 --> 00:04:45,929 is what captures that. 54 00:04:46,169 --> 00:04:52,009 And it's actually what I'm hearing a lot of people being concerned about at the conference. 55 00:04:52,009 --> 00:05:04,249 You have people who have been at the business for a long time and have a lot of knowledge, and then people who are new to the business and don't have that knowledge, and what's going to happen when all of that talent eventually does retire? 56 00:05:05,129 --> 00:05:10,569 There's a lot of promise in this approach as a way to capture 57 00:05:11,529 --> 00:05:16,089 the knowledge and strategies and principles that are really unspoken. 58 00:05:16,969 --> 00:05:18,729 So we're going to talk about that. 59 00:05:18,729 --> 00:05:20,249 We're going to talk about the normal go-to. 60 00:05:20,729 --> 00:05:27,849 Normal go-to with using large language models and they don't fit for you is to do some fine-tuning. 61 00:05:28,009 --> 00:05:29,089 There's some pitfalls to that. 62 00:05:29,089 --> 00:05:31,049 And we're going to talk about what to do instead. 63 00:05:32,009 --> 00:05:33,689 And we welcome questions throughout. 64 00:05:33,689 --> 00:05:35,529 So just feel free to raise your hand. 65 00:05:35,569 --> 00:05:38,049 And we'll also be around for questions afterwards. 66 00:05:39,769 --> 00:05:40,249 So 67 00:05:40,729 --> 00:05:42,329 Most Gen. 68 00:05:42,329 --> 00:05:46,329 AI projects fail due to lack of trust, not capability. 69 00:05:47,289 --> 00:05:56,169 JC was talking about there's average intelligence, there's great human intelligence, and then there's the best AI intelligence. 70 00:05:56,209 --> 00:06:03,649 So we don't need more intelligent models to do a lot of the use cases. 71 00:06:03,929 --> 00:06:05,209 that we're trying to do right now. 72 00:06:06,009 --> 00:06:07,289 But there are some impediments. 73 00:06:07,689 --> 00:06:09,529 One of them is trust. 74 00:06:09,529 --> 00:06:11,609 They say AI moves at the speed of trust. 75 00:06:12,969 --> 00:06:26,489 Does anybody want to offer what is being done within your organization to get past, to start to build trust with AI and actually have it be something that is more usable? 76 00:06:29,369 --> 00:06:29,769 Yes. 77 00:06:41,449 --> 00:06:42,329 Great, yep. 78 00:06:42,329 --> 00:06:51,449 Sharing what's worked, having that be kind of that teaching tool so people know that there are some successes out there and where it doesn't fit. 79 00:06:52,009 --> 00:06:52,569 Anybody else? 80 00:06:52,569 --> 00:06:52,729 Yes. 81 00:06:52,729 --> 00:07:09,729 Do you have a lot of read-only access for AI so that if they go to create ERP or any database, the user can see what to do with that data, interpret it, create visuals for that data without risking any harm? 82 00:07:11,369 --> 00:07:12,489 Yeah, I like that. 83 00:07:12,849 --> 00:07:14,569 People love that a lot. 84 00:07:14,569 --> 00:07:25,049 Read-only access that prevents the AI from going off the rails and deleting databases, but still seeing what the capabilities are while you're building trust. 85 00:07:25,209 --> 00:07:25,689 Awesome. 86 00:07:26,409 --> 00:07:27,529 How about one last one? 87 00:07:28,009 --> 00:07:28,409 Yes. 88 00:07:35,399 --> 00:07:42,999 So I guess we're more in the cautionary stage where it's a lot of things that recommend it's not. 89 00:07:45,329 --> 00:07:46,649 You can't distinct the answer. 90 00:07:46,649 --> 00:07:50,009 You have to know yourself and then trust it more as it gets better. 91 00:07:51,369 --> 00:07:51,649 Great. 92 00:07:51,649 --> 00:07:55,249 We're fine tuning that actually as we go. 93 00:07:55,249 --> 00:07:55,529 Awesome. 94 00:07:55,529 --> 00:08:00,169 So not necessarily trusting the output and building up. 95 00:08:00,409 --> 00:08:02,969 Using it as a tool but not you can't just trust it. 96 00:08:04,009 --> 00:08:04,969 No blind trust. 97 00:08:05,129 --> 00:08:05,609 Yes. 98 00:08:06,089 --> 00:08:06,489 Great. 99 00:08:08,169 --> 00:08:12,489 There's one other thing that is coming out of industry and it's this idea that 100 00:08:14,169 --> 00:08:26,409 We need some way to leverage the experts in our company and figure out how do you measure in comparison to how an expert would do a particular task. 101 00:08:27,129 --> 00:08:37,289 And that measurement capability is actually the thing that is the foundation upon which what we're talking about today is built. 102 00:08:37,689 --> 00:08:38,809 So measuring, 103 00:08:40,089 --> 00:08:44,169 how the AI is performing in industry terms. 104 00:08:44,169 --> 00:08:47,049 It's called LLM evaluations. 105 00:08:47,529 --> 00:08:56,049 So here's some industry quotes that talk about really if you're not investing in evals, you're not really shipping, you're guessing. 106 00:08:56,049 --> 00:08:58,729 A lot of pilots now are coming out of the gate. 107 00:08:58,729 --> 00:09:07,129 Pilots, when you're just trying to prove something out, are coming with some sort of measurement ability. 108 00:09:07,849 --> 00:09:09,049 And we'll talk more about that. 109 00:09:11,289 --> 00:09:14,249 So that's sort of the foundation. 110 00:09:14,249 --> 00:09:24,249 Then the MIT State of AI and Business report says, here's the big problem, the big barrier that everybody is running into. 111 00:09:24,969 --> 00:09:31,849 And again, it's not model quality, it's that these AI systems that we're interacting with aren't learning. 112 00:09:32,249 --> 00:09:32,729 So 113 00:09:33,289 --> 00:09:49,809 If you have a team or group or division department using AI and we're working with it and teaching it, like, here's how we do some things, oh, this isn't really documented, but this is how we like to work, none of that is captured. 114 00:09:49,929 --> 00:09:54,809 So it literally is Groundhog's Day all over again every day. 115 00:09:54,969 --> 00:09:57,609 The AI is just going to keep giving you the same answer. 116 00:09:58,329 --> 00:09:59,609 And that's the big problem. 117 00:10:01,769 --> 00:10:02,329 So 118 00:10:02,609 --> 00:10:07,609 What big tech does to solve for that, you have feedback loops. 119 00:10:07,609 --> 00:10:09,129 You have the thumbs up, thumbs down. 120 00:10:09,289 --> 00:10:10,329 We've all seen these. 121 00:10:10,649 --> 00:10:12,649 Thumbs down, then you can give feedback. 122 00:10:13,449 --> 00:10:16,009 But who is that helping? 123 00:10:16,169 --> 00:10:19,809 It's not helping your team or your company. 124 00:10:19,809 --> 00:10:29,209 It's helping big tech make a better model that really, you know, it's improving things for everybody, but not necessarily for your particular group. 125 00:10:32,969 --> 00:10:48,649 So, I heard MIT talks about AI's really great for individual brainstorming, but stalls out in enterprise settings because they lack iterative learning, and I actually heard a great quote at lunch. 126 00:10:49,769 --> 00:10:56,489 that was someone saying, I'm not interested in AI that is just for an individual. 127 00:10:56,489 --> 00:11:01,209 I want to be thinking about AI that is for multiple people, that's for teams. 128 00:11:01,969 --> 00:11:04,729 And that's where a lot of companies are right now. 129 00:11:04,729 --> 00:11:07,769 They want to break out from the individual use. 130 00:11:08,409 --> 00:11:12,409 So MIT gave this big problem, but they actually didn't give us an answer. 131 00:11:12,729 --> 00:11:14,409 They just said, it's a problem. 132 00:11:14,689 --> 00:11:18,969 And we think the solution is that you need some sort of loop, learning loop. 133 00:11:19,689 --> 00:11:20,649 but they didn't tell us how. 134 00:11:20,649 --> 00:11:22,649 And that's what we're sharing today. 135 00:11:24,889 --> 00:11:27,609 Fine-tuning, that's the typical go-to. 136 00:11:28,009 --> 00:11:34,849 Fine-tuning requires thousands, sometimes hundreds of thousands of examples to do fine-tuning. 137 00:11:34,849 --> 00:11:40,009 When you do fine-tuning, you are changing the weights in the model. 138 00:11:40,089 --> 00:11:43,289 These big, giant, smart models are actually nothing more than a CSV file. 139 00:11:43,609 --> 00:11:45,289 So you're changing 140 00:11:46,649 --> 00:11:58,249 The numbers in the CSV, when you do that, you have higher standards when it comes to governance and how the model is governed within your organization. 141 00:11:59,209 --> 00:12:06,969 It's really difficult to change fine-tuned models that are the ones that we like using. 142 00:12:06,969 --> 00:12:10,889 We like using Anthropic and OpenAI. 143 00:12:12,889 --> 00:12:16,809 But fine-tuning is something that typically happens with open weight models. 144 00:12:18,249 --> 00:12:31,129 And then finally, when you go to fine-tuning, which really kind of changes the behavior of the model so it can adapt to your domain, it's a whole other tier of expenses. 145 00:12:31,289 --> 00:12:34,009 Like you're basically getting an ML ops, AI ops. 146 00:12:34,489 --> 00:12:39,689 It's a lot of expense to do that instead of paying pennies per interaction. 147 00:12:41,289 --> 00:12:45,449 Okay, so those are some of the downfalls of fine-tuning. 148 00:12:46,569 --> 00:13:04,409 Another interesting thing as we get into the solution is LLMs are very good at doing what you tell them, and a lot of the failures that we encounter come from not being specific enough, not removing ambiguity. 149 00:13:05,769 --> 00:13:09,849 So most prompt failures are actually knowledge gaps. 150 00:13:10,169 --> 00:13:15,609 where some of these principles or strategies, we just didn't have a way, we didn't know how to express them. 151 00:13:15,609 --> 00:13:20,169 And then the AI gives us a result, and we're like, well, that's kind of stupid. 152 00:13:22,289 --> 00:13:27,049 But the key point here is that we all know what a prompt is, of course. 153 00:13:28,889 --> 00:13:35,769 Probably a lot of people know that there's also something called a system prompt or a system message. 154 00:13:35,769 --> 00:13:38,409 It's like a prompt that we never see. 155 00:13:39,049 --> 00:13:50,729 behind the scenes that says, be kind to the person you're working with and don't be, here are some ethical guidelines and here are some rules to follow. 156 00:13:51,289 --> 00:14:06,569 So the key point, though, is that prompt changes, changing that wording just a little bit can have massive differences on how the model performs. 157 00:14:07,369 --> 00:14:10,569 And this used to be prompt engineering. 158 00:14:10,569 --> 00:14:11,489 People made fun of it. 159 00:14:11,529 --> 00:14:13,849 was going to be a $200,000 career, proofed engineering. 160 00:14:15,929 --> 00:14:26,729 But there's actually what we're showing today is a mathematical, a scientific way to change the words that are in every single application that we're using. 161 00:14:27,209 --> 00:14:36,649 Even if you're just using a team copilot, there's a way to give it a prompt that helps it work with your team better. 162 00:14:37,169 --> 00:14:40,409 and recognize how your business is working. 163 00:14:41,449 --> 00:14:51,609 These small prompt changes are a way to mathematically, scientifically, methodically change the prompt to get better outcomes for you. 164 00:14:53,369 --> 00:15:04,969 This is a guy who talks a lot about AI on the internet, and he says LLMs just, they don't come with instructions in the box. 165 00:15:05,529 --> 00:15:06,649 So that's kind of the thing. 166 00:15:06,649 --> 00:15:09,129 We're all figuring out what are the right instructions. 167 00:15:09,129 --> 00:15:21,289 There was a talk early this morning that was sharing out what a model is and what agents are and what skills are and what tools are. 168 00:15:21,289 --> 00:15:27,929 So there are all of these places where you give instructions to the model. 169 00:15:28,489 --> 00:15:31,609 And again, we're figuring that out together. 170 00:15:32,329 --> 00:15:33,129 So some of the 171 00:15:33,769 --> 00:15:36,169 kind of go-tos that we have fine-tuning. 172 00:15:36,409 --> 00:15:37,929 Okay, it's expensive. 173 00:15:38,489 --> 00:15:39,369 Here's the downfall. 174 00:15:39,369 --> 00:15:42,409 RAG, that's where you pull in all of your documents. 175 00:15:42,409 --> 00:15:44,889 Maybe you have existing SOPs. 176 00:15:45,289 --> 00:15:46,249 That works great. 177 00:15:46,729 --> 00:15:47,609 We love RAG. 178 00:15:48,729 --> 00:16:00,409 But it doesn't help you bring out the reasoning and the strategy that your team is using when you're doing work within your organization. 179 00:16:01,049 --> 00:16:02,489 Memory is really good. 180 00:16:04,249 --> 00:16:10,089 But we don't want memorization of just things that have been done in the past. 181 00:16:10,089 --> 00:16:12,729 We want to extract lessons from that. 182 00:16:13,129 --> 00:16:20,089 So all of these things that we want just don't exist yet until this method that we're going to show you. 183 00:16:21,289 --> 00:16:30,489 If anybody's working on skills, skills is just a text file that helps you guide your AI to do a particular thing. 184 00:16:31,329 --> 00:16:32,649 And that 185 00:16:33,929 --> 00:16:36,289 actually isn't scaling very well. 186 00:16:36,289 --> 00:16:40,089 Like a lot of people are trying to figure out how do you scale that beyond just one person? 187 00:16:40,089 --> 00:16:42,249 How do you scale that to multiple people, A-team? 188 00:16:43,609 --> 00:16:45,449 So those are the problems that we're trying to answer. 189 00:16:46,009 --> 00:16:47,689 That's what we are going to answer today. 190 00:16:48,329 --> 00:16:52,209 And this is the kind of what if we could? 191 00:16:52,209 --> 00:17:00,009 What if we had a way to do all of this continuous improvement, to have it be fully auditable? 192 00:17:00,009 --> 00:17:01,689 Because instead of changing 193 00:17:02,569 --> 00:17:03,529 Model weights. 194 00:17:03,729 --> 00:17:07,009 these numbers in a giant CSV file, we're changing the instructions. 195 00:17:07,009 --> 00:17:10,329 And you can see, oh, the model is performing better because of these words. 196 00:17:10,969 --> 00:17:15,329 I can get that, like an auditor can look at that and say, I understand what that's doing. 197 00:17:16,649 --> 00:17:18,409 There's all of these great benefits. 198 00:17:19,689 --> 00:17:29,449 And they just so happen to get you similar to sometimes even better performance gains when compared to fine-tuning, the really expensive thing. 199 00:17:30,569 --> 00:17:32,409 I'm not going to hold you in suspense further. 200 00:17:32,409 --> 00:17:44,969 This is the library that came out that really does this magic thing that kind of codifies what the author at Stanford released. 201 00:17:45,129 --> 00:17:54,889 It allows you to take text and take some feedback on why that text was good or bad and make better text. 202 00:17:54,889 --> 00:17:56,969 It works on text. 203 00:17:57,369 --> 00:17:58,889 So it's super simple stuff. 204 00:17:59,769 --> 00:18:01,769 It can actually be done with a spreadsheet. 205 00:18:03,209 --> 00:18:11,249 Take some really bad interactions with your AI, talk to some subject matter experts, and say, why was it bad? 206 00:18:11,249 --> 00:18:13,529 What would an expert have said? 207 00:18:15,049 --> 00:18:16,489 And then run it through this. 208 00:18:17,169 --> 00:18:22,729 And what's interesting is you're not getting the hard-coded answers embedded into this text. 209 00:18:23,209 --> 00:18:33,849 you're getting higher-level principles and strategies extracted for you that can apply to your entire domain, examples not yet seen. 210 00:18:35,369 --> 00:18:37,849 Okay, so we're getting to an example. 211 00:18:38,249 --> 00:18:39,609 We're going to do a chat thing. 212 00:18:41,529 --> 00:18:45,609 Yeah, AI is not all about chat, but it's an example that we all recognize. 213 00:18:46,009 --> 00:18:49,769 So we ask a question, we get an answer, we give a thumbs up, thumbs down. 214 00:18:49,769 --> 00:18:52,649 We say it was a thumbs down because 215 00:18:53,689 --> 00:19:00,009 Someone's asking about photosynthesis because you didn't talk about light absorption details. 216 00:19:00,249 --> 00:19:03,449 So to be more precise, photosynthesis is a two-step blah, blah, blah. 217 00:19:04,649 --> 00:19:15,289 So we have a feedback loop for where the humans or maybe even smarter, super expensive models are saying, here's a better answer. 218 00:19:16,089 --> 00:19:20,169 We're applying that to food. 219 00:19:22,169 --> 00:19:29,729 We have lots of insurance and retail and energy and defense and med tech examples, but we're applying it to food. 220 00:19:29,729 --> 00:19:30,409 There's a rating. 221 00:19:32,489 --> 00:19:42,489 There's a rating system called Michelin Star that if you're traveling, these are like, hey, these are some useful places to go for an interesting experience. 222 00:19:42,809 --> 00:19:43,929 So it's like fancy food. 223 00:19:44,489 --> 00:19:45,529 But that's what we want to do. 224 00:19:46,009 --> 00:19:47,209 We have an example. 225 00:19:47,289 --> 00:19:48,409 It's running in code. 226 00:19:49,249 --> 00:19:53,929 You can stop by our booth and we'll show you the code if you want to see the details of it. 227 00:19:55,369 --> 00:20:02,329 But we're saying, OK, how would you make an omelet? 228 00:20:03,089 --> 00:20:08,649 And then how would a Michelin fancy chef make an omelet? 229 00:20:09,129 --> 00:20:15,849 So it's that difference between how anybody would do it and then how an expert would do it. 230 00:20:16,529 --> 00:20:17,929 And what are we doing behind the scenes? 231 00:20:17,929 --> 00:20:28,409 We're basically taking, we're asking Michelin-trained chefs, what is the really great answer of how you make an omelet? 232 00:20:28,609 --> 00:20:33,849 What are all the things that you need to consider, temperature and all the things, seasoning? 233 00:20:34,889 --> 00:20:45,769 And we use AI to say, okay, the answer that was given, this is a simplified example, but the answer that was given, how many of 234 00:20:46,089 --> 00:20:54,809 The bullet points or statements made by the expert chef were represented in the answer that the AI gave you. 235 00:20:55,529 --> 00:20:57,049 Maybe one out of five. 236 00:20:57,369 --> 00:21:12,329 So then we tell the AI why it was wrong, and we do that for 30, 40, 50 cycles, and run it through the system, and give it a new system prompt. 237 00:21:14,009 --> 00:21:20,889 And you get the same kind of performance gains that you get from the really super expensive reinforcement learning. 238 00:21:22,249 --> 00:21:23,529 Okay, that was a lot of detail. 239 00:21:25,129 --> 00:21:25,529 Food. 240 00:21:25,689 --> 00:21:26,449 We're talking about food. 241 00:21:26,449 --> 00:21:27,729 I'm going to hand it over to Matt. 242 00:21:27,729 --> 00:21:29,609 No, thank you. 243 00:21:30,249 --> 00:21:34,409 So as Matt had alluded to, we are talking food. 244 00:21:34,889 --> 00:21:41,449 And one of the challenges that we face in organizations is sometimes we don't know the question that we should be asking. 245 00:21:42,009 --> 00:21:45,529 So in this example, the question isn't really great. 246 00:21:45,529 --> 00:21:46,969 It's how do I make a roast? 247 00:21:47,209 --> 00:21:49,129 That's leaving out a lot of variables. 248 00:21:49,369 --> 00:21:51,289 Are there any home chefs in the room? 249 00:21:52,409 --> 00:21:56,369 So you may say this is a bad question because I don't know how big the roast is. 250 00:21:56,369 --> 00:21:57,929 I don't know what type of meat it is. 251 00:21:58,169 --> 00:22:01,689 I don't know what technique you're using. 252 00:22:02,169 --> 00:22:04,969 All of these lead to a very generic answer. 253 00:22:05,289 --> 00:22:08,169 Hey, let's flip every 30 to 45 minutes. 254 00:22:08,809 --> 00:22:10,169 not super helpful. 255 00:22:10,569 --> 00:22:17,049 And so we give a thumbs down and we explain that every 30 to 45 minutes, it's just actively bad advice. 256 00:22:17,209 --> 00:22:20,329 We need to provide better context. 257 00:22:20,729 --> 00:22:24,409 And 20 to 30 minutes per pound, well, that's a blunt instrument. 258 00:22:24,409 --> 00:22:27,929 We don't actually know without more of those details. 259 00:22:28,569 --> 00:22:36,969 So we give that feedback and we let our subject matter experts actually provide what is that ideal answer. 260 00:22:39,609 --> 00:22:44,569 We start with something very simple, like that answer had something like this in the prompt. 261 00:22:44,569 --> 00:22:52,889 It's like, hey, answer questions accurately, use any context that you have, but it doesn't even know that it's supposed to be acting like a Michelin star chef. 262 00:22:54,849 --> 00:22:57,689 When we are done with the process, this goes on and on. 263 00:22:57,689 --> 00:23:01,449 If you want to read the full ending prompt, feel free to scan that. 264 00:23:01,689 --> 00:23:09,609 But the interesting thing, as Matt had alluded to, is nowhere in the final result does the word roast even appear. 265 00:23:10,089 --> 00:23:15,169 Instead, we're talking about the scientific qualities of the answer. 266 00:23:15,169 --> 00:23:20,329 We're talking about the Mylar reaction, the browning that you get on your meat when you're cooking it. 267 00:23:20,889 --> 00:23:24,969 It's talking about collagen structures, what oils to use in what cases. 268 00:23:25,609 --> 00:23:30,809 Said simply, it learned the principles, not the answers to the questions that we are asking. 269 00:23:32,449 --> 00:23:34,689 And we can see a very strong result. 270 00:23:34,889 --> 00:23:36,969 On the left here, how do you make a roast? 271 00:23:37,129 --> 00:23:38,089 This is the original. 272 00:23:38,329 --> 00:23:39,529 Same bad answer. 273 00:23:39,849 --> 00:23:47,849 But with that updated prompt, it starts out by telling you, need to choose the right cut, asking you immediately, beef, pork, or lamb? 274 00:23:48,329 --> 00:23:50,489 tells you how to prepare the meat, how to season it. 275 00:23:50,489 --> 00:23:53,969 Did you have something else? 276 00:23:53,969 --> 00:23:55,049 I was just going to add one thing. 277 00:23:55,529 --> 00:24:06,809 By the way, it's been proven that if you say act like a super experienced data scientist or act like a super experienced Michelin star chef, that doesn't work. 278 00:24:07,369 --> 00:24:08,889 That kind of prompting does not work. 279 00:24:09,529 --> 00:24:13,689 It did work a few years ago, but the models have grown since then. 280 00:24:15,529 --> 00:24:17,049 And why is this a problem? 281 00:24:17,049 --> 00:24:19,689 Why can't we just have somebody on our team sit down and write it? 282 00:24:20,249 --> 00:24:28,729 Well, JC had mentioned in the keynote that it is somebody's job at Anthropic to write the system prompt or the sole document. 283 00:24:29,689 --> 00:24:31,529 This document, they've been leaked. 284 00:24:31,529 --> 00:24:36,089 Anthropic does also release these themselves after some delay. 285 00:24:36,809 --> 00:24:41,769 About 24, 25,000 words and many, many lines long. 286 00:24:42,689 --> 00:24:50,089 I have never been part of a team that can actually justify spending that much time on one document every single cycle. 287 00:24:51,449 --> 00:24:58,569 So that is where JEPA, the parent library to optimize anything that Matt had introduced, comes in. 288 00:24:59,129 --> 00:25:07,609 This is a result of a research paper out of Stanford saying that reflective prompt evolution can outperform reinforcement learning. 289 00:25:07,929 --> 00:25:08,729 Essentially, 290 00:25:09,209 --> 00:25:12,969 Give the AI a signal, and it's a better prompt engineer than you or I. 291 00:25:15,769 --> 00:25:24,569 And what it does is it helps us extract the principles, practices, strategies, and techniques, specifically not rote memorization of what the answer should have been. 292 00:25:25,049 --> 00:25:32,569 This is important because, yeah, if we just say, when asked how to cook a roast, respond with this, it will always be correct. 293 00:25:33,289 --> 00:25:37,449 But it does not generalize to every other recipe that I may want to attempt. 294 00:25:40,249 --> 00:25:42,329 So how might we leverage this under the hood? 295 00:25:42,329 --> 00:25:44,569 How does optimize anything really work? 296 00:25:44,889 --> 00:25:47,129 Well, it's doing something very similar to this. 297 00:25:47,129 --> 00:25:49,769 I went into ChatGPT, typed this question. 298 00:25:49,769 --> 00:25:54,729 It was very helpful and gave me the same, the right formatting on the output. 299 00:25:55,209 --> 00:25:58,529 But essentially, we need to collect those weird examples. 300 00:25:58,529 --> 00:25:59,609 Where does it fail? 301 00:26:00,489 --> 00:26:03,529 This is that thumbs down, by the way, from our subject matter experts. 302 00:26:04,489 --> 00:26:06,089 And then we request improvement. 303 00:26:06,089 --> 00:26:14,969 We go in and we tell our AI, I used this prompt, this system prompt, this input, the user's question, and it gave me this weird answer. 304 00:26:15,449 --> 00:26:17,489 And it was weird for these reasons. 305 00:26:17,489 --> 00:26:22,169 And then it gives us a new prompt, and we can try again with our subject matter experts. 306 00:26:23,129 --> 00:26:28,169 Now, this doesn't scale very well, but it's a really good starting point. 307 00:26:29,369 --> 00:26:31,929 Now, what do we do if that doesn't scale? 308 00:26:32,409 --> 00:26:33,449 We can use 309 00:26:34,569 --> 00:26:38,409 programming, development, to automate this process. 310 00:26:38,969 --> 00:26:45,689 We can use that feedback and let the system actually write its own feedback and say it's correct for these reasons. 311 00:26:45,689 --> 00:26:54,809 It included the right oil, it included the information about the Mylar reaction, but it missed information about the internal temperature. 312 00:26:56,569 --> 00:27:03,849 The difference between a traditional optimization in machine learning, also known as gradient descent, and JEPA is that 313 00:27:04,329 --> 00:27:09,129 This is the only signal that a traditional system will get is that number, 0 to 1. 314 00:27:09,609 --> 00:27:14,649 Not a whole lot of understanding of where did we go right and where did we go wrong. 315 00:27:15,929 --> 00:27:18,969 So that natural language feedback is invaluable. 316 00:27:20,329 --> 00:27:22,969 Let's take a look at a little bit of a visual here. 317 00:27:23,769 --> 00:27:30,409 In machine learning, large language models, this is actually what the inside of the brain kind of looks like conceptually. 318 00:27:30,969 --> 00:27:32,209 We have all of these hills. 319 00:27:32,209 --> 00:27:33,849 These are all of the expert 320 00:27:34,329 --> 00:27:35,129 topics. 321 00:27:35,529 --> 00:27:39,129 And that red ball there, that is what a traditional process will do. 322 00:27:39,449 --> 00:27:46,969 It starts at some point in the map, and it starts trying to climb the hills around it, getting to the highest point on that plane. 323 00:27:47,689 --> 00:27:54,409 The problem is it climbs to not quite the highest hill, but to something that is, oh, middle of the road. 324 00:27:55,049 --> 00:27:58,889 But because of the way that technology works, it gets stuck. 325 00:27:59,689 --> 00:28:03,649 Meanwhile, JEPA, the blue ball, or the green ball, sorry, 326 00:28:04,569 --> 00:28:18,249 is actually jumping around because it's given human feedback, and it's able to see, oh, I need to jump over here, and eventually it found that peak much faster and with hundreds, not thousands, of examples. 327 00:28:18,969 --> 00:28:26,649 You can actually run this with as few as 10 examples, but in our example, I think we had 200 question and answer pairs from experts. 328 00:28:28,329 --> 00:28:31,129 So the question that I have for everyone is, 329 00:28:31,529 --> 00:28:32,969 We're not optimizing prompts. 330 00:28:32,969 --> 00:28:34,249 We're optimizing text. 331 00:28:34,409 --> 00:28:37,769 Where else might we see text in our AI applications? 332 00:28:38,649 --> 00:28:39,369 Any examples? 333 00:28:44,249 --> 00:28:46,089 Earlier we had talked about skills. 334 00:28:46,249 --> 00:28:50,489 There's also MCP servers, tool routing logic. 335 00:28:51,049 --> 00:28:52,969 All of these are possible. 336 00:28:53,209 --> 00:28:56,329 The 2 that this demo focuses on are the system prompts. 337 00:28:56,649 --> 00:28:59,129 and what is called an LLM as a judge. 338 00:28:59,529 --> 00:29:03,129 This is what powers the evaluations that Matt had talked about. 339 00:29:03,369 --> 00:29:13,929 This is a stand-in for our users so that we can test 10s to hundreds of different prompts along the way without driving our subject matter experts up a wall. 340 00:29:16,169 --> 00:29:16,569 So 341 00:29:17,529 --> 00:29:20,329 We are almost through all of the math heavy, I promise. 342 00:29:20,729 --> 00:29:24,889 But this was too cool not to show, so I wanted to bring this to light. 343 00:29:25,289 --> 00:29:30,249 This is why JEPA is so effective compared to traditional methods. 344 00:29:31,049 --> 00:29:35,569 That 0 circle at the top, that is where we start. 345 00:29:35,569 --> 00:29:38,089 That's the initial, how do I make a roast question. 346 00:29:38,089 --> 00:29:39,369 It didn't do very good. 347 00:29:40,089 --> 00:29:45,849 And it tried five different methods to, or five different prompts to improve. 348 00:29:46,489 --> 00:29:49,209 And #5 was the winner of that generation. 349 00:29:49,689 --> 00:29:53,609 In a traditional world, we would have thrown out one through 4. 350 00:29:54,089 --> 00:30:06,729 But JEPA allows us to continuously explore that space, eventually landing all the way over here on the left on child 12 that was a descendant of 1, which performed quite a bit worse than #5. 351 00:30:07,129 --> 00:30:12,009 It allows us to more effectively explore the expertise of our language models. 352 00:30:12,009 --> 00:30:14,169 And a really cool thing about 353 00:30:15,689 --> 00:30:25,769 This optimize anything is that the library now spits out a graph like this, and you can hover over each one of those nodes and see how the prompt has evolved. 354 00:30:25,769 --> 00:30:29,689 So hover over node 0, it's you're a helpful agent, try not to be rude. 355 00:30:30,329 --> 00:30:40,129 And then hover over node 12, and it has all the stuff in there about you need to pay attention to flavor and how you lock it in and the chemistry of cooking. 356 00:30:40,129 --> 00:30:43,129 Yeah, that's a great call out, Matt. 357 00:30:45,369 --> 00:30:50,249 So really, the difference said simply is the old way is this version feels better. 358 00:30:50,249 --> 00:30:51,369 We got a higher number. 359 00:30:51,769 --> 00:30:58,089 The new one says we know that it performs in these categories very well for our evaluations. 360 00:30:58,489 --> 00:31:00,649 We no longer are asking, is it good enough? 361 00:31:00,889 --> 00:31:08,809 We can now constantly say this prompt sits on the, how do we say it, the efficient frontier of our evaluations. 362 00:31:09,049 --> 00:31:12,569 It said simply aligns with our subject matter experts. 363 00:31:13,849 --> 00:31:16,329 So what actually changes when we do this? 364 00:31:16,649 --> 00:31:18,649 We're changing the instructions and logic. 365 00:31:18,809 --> 00:31:22,169 Matt had mentioned this is an auditor's dream. 366 00:31:22,569 --> 00:31:26,169 No longer are we looking at why is this a.5 instead of a.4? 367 00:31:26,489 --> 00:31:31,449 We're looking at make sure to remember the Mylar reaction. 368 00:31:31,849 --> 00:31:38,809 Know to use avocado oil in these scenarios, olive oil in these scenarios, and sunflower oil here. 369 00:31:39,929 --> 00:31:42,089 We aren't locked into specific models. 370 00:31:42,089 --> 00:31:50,489 We can continue to use the Clauds and the GPTs of the world, but we still have the flexibility to use those open weight models if we choose. 371 00:31:51,289 --> 00:31:57,529 And because it's just instructions, it means that undoing these changes takes minutes, not days or weeks. 372 00:32:01,529 --> 00:32:08,809 Before, our code looked something like this, and our average score, it was getting about 67% of the answer correct. 373 00:32:09,209 --> 00:32:18,249 And if we looked for strict accuracy, meaning it hit every bullet point that an expert cared about, we only got 35% of the questions correct. 374 00:32:19,449 --> 00:32:24,249 So afterwards, we saw task-specific information. 375 00:32:24,249 --> 00:32:26,329 It very clearly described the inputs. 376 00:32:26,729 --> 00:32:29,209 It described what output it's expecting. 377 00:32:29,689 --> 00:32:31,209 Lean into food science. 378 00:32:32,009 --> 00:32:35,209 It added the domain-specific knowledge, the food chemistry. 379 00:32:35,849 --> 00:32:43,609 And most importantly, it defines strategies, understanding the difference between different cooking methods and the trade-offs. 380 00:32:45,369 --> 00:32:47,449 And the results speak for themselves. 381 00:32:47,689 --> 00:32:50,409 We didn't change the data available to the system. 382 00:32:50,569 --> 00:32:52,249 We only changed the instructions. 383 00:32:53,089 --> 00:33:00,489 And we ended up with a 9.7% improvement in the average score and an 8.8 in improvement in strict accuracy. 384 00:33:00,729 --> 00:33:04,249 Again, this is just from listening to our subject matter experts. 385 00:33:06,649 --> 00:33:13,609 So a couple of other in the industry at scale, we see this is an example from Shopify. 386 00:33:14,089 --> 00:33:20,169 Shopify runs one of the largest e-commerce platforms in the world, probably only second to like Amazon. 387 00:33:20,969 --> 00:33:27,129 And they were running a very expensive system, analyzing every storefront. 388 00:33:27,769 --> 00:33:32,329 They were spending millions of dollars a year, and they covered 13% of stores. 389 00:33:32,889 --> 00:33:44,889 Not very great they used JEPA to actually train a smaller model to be more effective, and now they can cover 100% of shops seventy-five times cheaper. 390 00:33:45,769 --> 00:33:53,369 They got over five times the ability, and they spent seventy-five times less. 391 00:33:54,769 --> 00:34:06,009 Another example from Dropbox, if you don't know when you search for a file in Dropbox, your search and the files are actually going to AI and saying, hey, does this file and description match this search term? 392 00:34:06,729 --> 00:34:22,409 And again, Dropbox used JEPA to use, again, a smaller model and lower their adaptation to changing needs from their users from weeks to days, all just from listening to feedback. 393 00:34:24,409 --> 00:34:26,009 This is a bit more local. 394 00:34:26,089 --> 00:34:34,489 I'm currently working with a Fortune 500 client, and they're having AI write queries against their data lake. 395 00:34:34,489 --> 00:34:36,249 Think Databricks or Power BI. 396 00:34:36,249 --> 00:34:43,049 When we started out, the AI knew nothing about that environment, and it was only scoring a 58%. 397 00:34:43,049 --> 00:34:50,409 In under an hour and less than $5 worth of AI usage, we got all the way up to 89%. 398 00:34:50,969 --> 00:34:55,689 Again, all just leveraging existing knowledge from the team, saying yes or no. 399 00:34:57,849 --> 00:35:00,769 So again, let's go ahead through what changed. 400 00:35:00,889 --> 00:35:03,289 We went domain-specific rather than general. 401 00:35:03,609 --> 00:35:06,329 We kept everything task-specific. 402 00:35:06,489 --> 00:35:10,969 We learned strategies and prescribed what output we actually cared about. 403 00:35:11,529 --> 00:35:12,569 All auditable. 404 00:35:12,569 --> 00:35:14,329 Our data governance folks love it. 405 00:35:16,569 --> 00:35:18,089 So let's flip the script. 406 00:35:18,329 --> 00:35:32,649 Instead of giving the thumbs up and thumbs down information to ChatGPT, to Claude, to Copilot, let's bring that back internally and improve our own products, creating the competitive advantage instead of just rising with the tide. 407 00:35:34,889 --> 00:35:37,609 I'll leave you all with an architecture overview. 408 00:35:37,769 --> 00:35:40,409 This outlines what we've been talking about today. 409 00:35:40,969 --> 00:35:43,689 That chat has a thumbs up and thumbs down. 410 00:35:43,689 --> 00:35:45,369 Our users can provide feedback. 411 00:35:45,769 --> 00:35:48,649 And then all of that goes into this optimization pipeline. 412 00:35:48,889 --> 00:35:54,729 We store that in a tool, an open source tool called Phoenix, so that the data never leaves our customer's environment. 413 00:35:55,449 --> 00:35:58,249 And then that is used to train up a user judge. 414 00:35:58,569 --> 00:36:06,329 And the user judge, along with the user feedback, allows us to optimize and say, yes or no, we are actually improving. 415 00:36:06,889 --> 00:36:16,249 This allows us to run on a weekly or monthly basis, depending on the amount of feedback we've gotten, and continuously improve with how the team uses the tool. 416 00:36:17,689 --> 00:36:22,889 I'll pass it back over to Matt to close this out, and then we'll be ready for some questions. 417 00:36:23,849 --> 00:36:25,929 So yeah, this is second to last slide. 418 00:36:26,409 --> 00:36:28,649 This is if you want to use this tomorrow. 419 00:36:29,769 --> 00:36:30,569 Here's what you can do. 420 00:36:31,369 --> 00:36:34,009 Yes, you need access to a developer who will 421 00:36:34,569 --> 00:36:36,649 pull down that optimize anything library. 422 00:36:37,209 --> 00:36:40,329 But all you need to give them is a spreadsheet. 423 00:36:41,049 --> 00:36:47,769 So a spreadsheet is 10 to 30 interactions with an AI system. 424 00:36:48,569 --> 00:36:53,849 And then you sit that down in front of the expert and you say, where did this go wrong? 425 00:36:54,969 --> 00:36:57,369 And they write down, okay, here's where they went. 426 00:36:57,609 --> 00:36:58,889 This was like completely off. 427 00:36:58,889 --> 00:37:02,009 This they got, this was actually a good part of the answer. 428 00:37:02,889 --> 00:37:05,849 But then you get the expert feedback in there. 429 00:37:06,489 --> 00:37:08,089 You feed it to optimize anything. 430 00:37:08,809 --> 00:37:10,889 You spit out the text. 431 00:37:11,209 --> 00:37:12,409 That's all it is text. 432 00:37:12,409 --> 00:37:14,889 It doesn't matter how you're building your Gen. 433 00:37:14,889 --> 00:37:15,289 AI app. 434 00:37:15,369 --> 00:37:21,769 And there is some place, no matter what you're using, Office Copilot to custom Gen. 435 00:37:21,769 --> 00:37:31,449 AI, there's some place where you can drop in this text and get big improvements, again, without the expense of fine tuning. 436 00:37:31,449 --> 00:37:31,769 So 437 00:37:33,769 --> 00:37:36,009 We're obviously super excited about it. 438 00:37:36,329 --> 00:37:40,969 Hopefully some of that enthusiasm rubs off. 439 00:37:41,129 --> 00:37:58,849 And let's see, we have one other question, or one other thing that there's like resources that we have in QR codes that are at our booth where the sponsor area is, and Ben and I will be over there to 440 00:37:59,449 --> 00:38:05,209 take any deep dives for anybody who wants to dig into code or more specifics if you'd like to. 441 00:38:06,249 --> 00:38:11,049 But love to hear what questions you have or where we can provide some clarity. 442 00:38:11,529 --> 00:38:12,169 Yes, Adam. 443 00:38:12,169 --> 00:38:28,329 I'm just curious to understand how this compares to the Andre Carpathy approach to self-improvement and if that's been played into this model or this way of approaching improvement. 444 00:38:30,169 --> 00:38:34,969 the question for the recording, sorry everyone, that mic does not go through the recording, so I'll be repeating. 445 00:38:35,769 --> 00:38:43,689 The question was, how does this compare to the Andrea Kaparthy auto research that was unveiled, what was it, maybe a month ago? 446 00:38:44,569 --> 00:38:47,769 This is, they're very similar in concept. 447 00:38:48,089 --> 00:38:52,009 JEPA is a year, year and a half old at this point. 448 00:38:52,329 --> 00:38:55,129 So we're auto research is really focused on 449 00:38:56,249 --> 00:39:03,129 Architecture of training models this is very much optimizing text they can be one and the same. 450 00:39:04,249 --> 00:39:08,569 Because again, the code to train a model is also text. 451 00:39:08,969 --> 00:39:13,369 So I would say that both are feasible and show the same promise. 452 00:39:15,849 --> 00:39:21,449 Yeah, and one of the QR codes is an Andrei Karpathy post because all things lead back to Andrei Karpathy. 453 00:39:21,769 --> 00:39:22,009 Yes. 454 00:39:32,129 --> 00:39:32,689 Adam again. 455 00:39:35,689 --> 00:39:36,849 Why did you decide Phoenix? 456 00:39:36,929 --> 00:39:46,249 And what's the significance behind the Phoenix portion of your database? 457 00:39:46,289 --> 00:39:52,249 The observability in AI is really important. 458 00:39:52,249 --> 00:40:00,649 So you need some ability to kind of log the traces or the interactions, the turns between the user and the AI. 459 00:40:01,609 --> 00:40:02,089 And 460 00:40:03,129 --> 00:40:12,329 Basically what we're doing with that product, and it's one of the products that Ben, the open source products that Ben supports and helps. 461 00:40:13,529 --> 00:40:22,329 But basically what we do from there is we pick, we kind of click through a bunch of things where we said, well, these are really great examples or these are really bad examples. 462 00:40:22,889 --> 00:40:25,609 And we do the classic data science thing. 463 00:40:25,609 --> 00:40:27,209 We turn it into a data set. 464 00:40:27,609 --> 00:40:29,049 We carve, we tag. 465 00:40:30,009 --> 00:40:31,689 80% of it for training data. 466 00:40:31,689 --> 00:40:35,929 We hold out 20% for validation data. 467 00:40:36,289 --> 00:40:42,329 And then we point optimize anything to that data set and have it give a new prompt. 468 00:40:42,849 --> 00:40:45,769 And we set the new prompt in Phoenix. 469 00:40:45,769 --> 00:40:47,209 Phoenix also stores prompts. 470 00:40:47,329 --> 00:40:50,969 And then our app just automatically pulls in the new prompt text. 471 00:40:52,329 --> 00:40:54,009 That was one way of explaining it. 472 00:40:54,249 --> 00:40:59,609 Yeah, I'll go ahead and echo what you said, Matt, is absolutely, it all holds true. 473 00:41:00,889 --> 00:41:05,609 A little bit of a different reason of why we choose Phoenix is it is open source. 474 00:41:05,609 --> 00:41:08,649 We don't have to worry about the data residency problem. 475 00:41:09,049 --> 00:41:17,849 If we have clients that are all on-prem or all in their own cloud, it makes it really easy for us to adhere to those. 476 00:41:20,489 --> 00:41:30,729 to those desires of keeping everything private, not sending off our very valuable LLM interactions and really company data to a third party. 477 00:41:30,889 --> 00:41:32,649 We're able to stay in control of that. 478 00:41:33,129 --> 00:41:40,249 And then on top of that, the feature set of Phoenix just all meshes very, very well with a system like Optimize Anything. 479 00:41:45,969 --> 00:41:47,129 Anyone other than Adam? 480 00:41:54,329 --> 00:41:54,689 He is. 481 00:41:54,689 --> 00:42:07,689 I guess I could ask what I've been asking pretty much everybody, but I teach engineering transfer courses for students that are going on to a four-year program from DNAC. 482 00:42:07,689 --> 00:42:21,369 And so I just, most of what I'm interested in here today, and he also teaches at DNAC, but what are the main tools that you would say is important that we make sure our students understand before they go out into the workforce? 483 00:42:22,129 --> 00:42:26,289 Two, three, four years from now, Daniel. 484 00:42:26,489 --> 00:42:31,769 Yeah, that's an interesting question, and I think is unfortunately a little bit... 485 00:42:34,289 --> 00:42:37,129 dependent on what path they choose to take. 486 00:42:37,449 --> 00:42:46,089 I would give different advice to somebody looking to become a data scientist, to somebody being a data engineer, different advice to a software engineer. 487 00:42:46,409 --> 00:42:58,489 So I think the general advice that I would give to anyone going from a two-year to a four-year degree like DMACC to Iowa State, as an example, would be remain curious, 488 00:42:59,129 --> 00:43:00,329 learn to learn. 489 00:43:00,649 --> 00:43:04,089 Don't get too hung up on one specific tool set. 490 00:43:04,489 --> 00:43:19,769 Because we've seen with the age of AI, things change so frequently that if we spend too much time making sure that this one tool set is perfect, we run the risk of that being out of date by the time they're out of school. 491 00:43:21,209 --> 00:43:25,489 We saw this back in the early 2010s with Hadoop. 492 00:43:25,849 --> 00:43:26,369 clusters. 493 00:43:26,369 --> 00:43:27,689 They were all the rage. 494 00:43:27,689 --> 00:43:28,649 Everyone had it. 495 00:43:28,889 --> 00:43:30,009 You need to go into Hadoop. 496 00:43:30,529 --> 00:43:34,809 And now I haven't worked with anyone that has a Hadoop cluster in a few years. 497 00:43:35,369 --> 00:43:39,609 So that is where I would be leaning is learn how to learn. 498 00:43:40,169 --> 00:43:41,769 Don't be loyal to any one tool. 499 00:43:42,569 --> 00:43:44,289 Understand that judgment point. 500 00:43:44,409 --> 00:43:45,689 Matt, do you have anything to add there? 501 00:43:47,049 --> 00:43:47,849 No, great answer. 502 00:43:54,089 --> 00:44:11,889 If I'm building a domain-specific AI tool for AEC workflows using like Revit API, where would you start with the self-improvement loop, talking with like the feedback on incorrect element detection, missed clashes, or something else? 503 00:44:13,609 --> 00:44:14,809 I'm sorry, I'm not familiar. 504 00:44:14,969 --> 00:44:17,609 Could you, what is an AEC environment there? 505 00:44:19,609 --> 00:44:22,969 Sorry, like construction industry. 506 00:44:26,089 --> 00:44:27,849 Matt, do you have any thoughts there? 507 00:44:32,569 --> 00:44:45,289 The incredible thing about this is that it really is the expert feedback and saying, okay, here's why a clash was missed. 508 00:44:45,769 --> 00:44:49,289 Here's what I know from my 30 years of experience. 509 00:44:50,169 --> 00:44:52,569 And doing that 30, 40, 510 00:44:53,129 --> 00:45:04,969 100, 200 times, or just setting up this loop that just every two weeks, it just pulls in anybody who gave any feedback and updates the prompt to make it better. 511 00:45:04,969 --> 00:45:11,289 So it doesn't matter what the strategy is, you're extracting those out. 512 00:45:13,929 --> 00:45:20,969 So it really becomes a domain-agnostic way to make domain-specific AI.