1 00:00:00,191 --> 00:00:02,751 All right, I'll go ahead and do a quick introduction. 2 00:00:02,751 --> 00:00:05,111 I'm Gail Masberg, and I'm Cirrus Marketing Manager. 3 00:00:05,111 --> 00:00:11,831 If you're new to this room, I get the pleasure to work for Cirrus in a little bit different lens. 4 00:00:11,831 --> 00:00:22,831 So I'm more on the internal side, so working on getting y'all to these events and helping prepare for these events, as well as working on website marketing PR for Cirrus. 5 00:00:22,991 --> 00:00:24,831 So welcome, that's me. 6 00:00:26,191 --> 00:00:28,551 get to introduce today coming off of lunch. 7 00:00:28,551 --> 00:00:30,671 Hopefully you all had a great energizing lunch. 8 00:00:31,791 --> 00:00:37,791 I get to introduce Mahul Abuva, a senior software engineer in data and AI at QCI. 9 00:00:38,031 --> 00:00:49,871 With more than 20 years of experience, Mahul has built and modernized enterprise systems using Microsoft Azure and advanced AI platforms, delivering scalable solutions across industries. 10 00:00:50,671 --> 00:00:55,951 He is also a recognized contributor to the developer community, author of SharePointFix.com. 11 00:00:56,431 --> 00:00:57,151 I've been there. 12 00:00:57,991 --> 00:01:03,071 a widely read technical blog and Microsoft-nominated Azure Developer Influencer. 13 00:01:03,471 --> 00:01:07,631 His work spans production-grade AI systems, including large-scale implementations. 14 00:01:08,271 --> 00:01:24,831 In today's session, Mahul will walk us through how to design and deploy enterprise-grade, read it all out, retrieval augmented generation, or RAG systems on Azure, sharing practical architecture insights, real-world examples, and best practices you can apply in your own organization. 15 00:01:25,471 --> 00:01:27,871 So please join me in welcoming Mehul. 16 00:01:29,711 --> 00:01:30,031 Thank you. 17 00:01:31,951 --> 00:01:34,911 Hello, Can you guys hear me out back there? 18 00:01:35,671 --> 00:01:35,871 Good? 19 00:01:36,031 --> 00:01:36,231 Okay. 20 00:01:36,831 --> 00:01:38,111 Let me know if I'm being too loud. 21 00:01:38,711 --> 00:01:39,791 I have that habit of being loud. 22 00:01:41,231 --> 00:01:44,191 So thank you, Gail, for introducing me. 23 00:01:44,431 --> 00:01:46,271 And this is Mehul Bhuva. 24 00:01:46,271 --> 00:01:50,111 I have 22 years of software development experience. 25 00:01:51,071 --> 00:02:07,871 I started with Microsoft technology stack as a full-stack app developer working on several Microsoft technologies, including VDASP, customizing SharePoint, C#,.NET, the classic world, the modern world,.NET Core, Angular, React, you name it. 26 00:02:08,031 --> 00:02:09,951 I've done all kinds of developments. 27 00:02:10,271 --> 00:02:15,391 But recently, I've stepped into the data and AI world. 28 00:02:15,951 --> 00:02:20,591 It's been three or four years now where I stepped in from an app world to a data world. 29 00:02:20,991 --> 00:02:25,711 and I can bring in a lot of experience and realize that I was able to contribute a lot there. 30 00:02:27,311 --> 00:02:29,791 So that's a background of my experience. 31 00:02:30,911 --> 00:02:32,231 Moving on to the next slides. 32 00:02:32,231 --> 00:02:35,391 Again, this Gail has already covered this part, so I'm not going to cover it. 33 00:02:35,871 --> 00:02:39,791 So the agenda for today, I don't want to bore you with the slide decks and slide shows. 34 00:02:39,791 --> 00:02:43,071 I know you guys are here for the demo, and I have 3 interactive demos. 35 00:02:44,111 --> 00:02:47,951 In fact, I build these two agents where I can show you the capabilities of the agent. 36 00:02:48,111 --> 00:02:50,991 And then the third demo is going to be we build the agent together. 37 00:02:52,351 --> 00:03:10,751 So the agenda is that we cover what is RAG and why RAG, RAG architecture, Azure AI Foundry, what Azure AI Foundry is all about, how to build on Azure AI Foundry, what are the building blocks, production design, how to design A production grade platform, 38 00:03:11,071 --> 00:03:14,591 for your or a framework for your rank-based scenarios. 39 00:03:15,711 --> 00:03:24,191 And a case study which I recently implemented at my client location, which is Corteva, I've been working with them for the last six years. 40 00:03:24,511 --> 00:03:31,951 We implemented an enterprise-wide chatbot for them that scales for up to 22,000 users across the globe. 41 00:03:33,631 --> 00:03:37,311 And so what are some of the best practices that I have learned from my experience? 42 00:03:38,431 --> 00:03:39,391 So moving on. 43 00:03:40,191 --> 00:03:40,911 So what is RAG? 44 00:03:40,911 --> 00:03:42,071 So what is RAG? 45 00:03:42,191 --> 00:03:43,391 Are you guys aware of RAG? 46 00:03:43,391 --> 00:03:44,671 First of all, let me ask you this. 47 00:03:45,951 --> 00:03:46,071 Okay. 48 00:03:46,591 --> 00:03:54,271 So pretty informed audience, retrieval, R stands for retrieval, A for augmented, and G for generation. 49 00:03:54,751 --> 00:03:57,871 So coming back here to RAG, right? 50 00:03:58,031 --> 00:04:02,511 So let's assume you have, and every company has this problem, right? 51 00:04:02,511 --> 00:04:05,551 You have an employee who joins on day one. 52 00:04:05,991 --> 00:04:06,271 right? 53 00:04:06,591 --> 00:04:09,631 He needs information right away to get started. 54 00:04:10,031 --> 00:04:11,751 It could be any kind of information. 55 00:04:11,751 --> 00:04:14,111 It could be an onboarding information, right? 56 00:04:14,351 --> 00:04:17,151 I need my, I need to sign up for a 401k plan. 57 00:04:17,151 --> 00:04:19,871 I need to sign up for a dental, medical, vision plans. 58 00:04:20,271 --> 00:04:21,871 How does the employee know all this? 59 00:04:22,111 --> 00:04:28,591 Secondly, if you have a plant operator at a site, he needs a checklist on day one, how he should learn the whole process. 60 00:04:28,831 --> 00:04:29,951 There are SOPs. 61 00:04:30,271 --> 00:04:33,791 There are PDFs, there are hundreds and hundreds of pages of documents that he has to go through. 62 00:04:34,151 --> 00:04:42,191 A majority of the companies, small, big, medium, doesn't matter, have large-scale, unstructured and semi-structured documents. 63 00:04:42,591 --> 00:04:48,191 They have lesser databases, but more documents sitting out there in a repository somewhere. 64 00:04:48,351 --> 00:04:50,511 It could be a OneDrive location. 65 00:04:51,071 --> 00:04:52,911 It could be a SharePoint site. 66 00:04:53,791 --> 00:04:56,831 or it could be a person's own physical machine. 67 00:04:56,991 --> 00:05:05,871 There are documents everywhere, and there are several versions of these documents, making it very difficult to find relevant information when you need it the most. 68 00:05:06,351 --> 00:05:07,151 That's the key here. 69 00:05:07,551 --> 00:05:12,991 So the employee who's starting on day one needs a lot of information on day one to get started on his job. 70 00:05:13,511 --> 00:05:21,231 Of course, there'll be onboarding, there'll be training, things like that, but if he's on the shop floor and he does not know what to do, then we have a bigger problem. 71 00:05:21,551 --> 00:05:27,391 because then the enterprise or the organization is not allowing him or enabling him to be productive from day one. 72 00:05:27,791 --> 00:05:30,591 That's why the whole concept of rack-based chatbot. 73 00:05:30,991 --> 00:05:39,311 So when I say this, it is very serious because we have seen the benefits of implementing a rack-based chatbot at Cortiva. 74 00:05:40,031 --> 00:05:43,551 So we have these hundreds of plant sites and operators across the globe. 75 00:05:43,951 --> 00:05:47,711 And when we provision, and they don't like to read documents, 76 00:05:48,071 --> 00:05:50,031 especially the documents, especially from Europe. 77 00:05:50,031 --> 00:05:52,351 They don't like to read documents that are in English. 78 00:05:52,591 --> 00:05:56,351 They like to go and translate them in their local language and then they can make sense out of it. 79 00:05:57,551 --> 00:06:10,911 So when you put the LLM, which is a large language model, in front of your own enterprise documents, it all of a sudden has more context because it is not looking at the world wide web. 80 00:06:11,231 --> 00:06:14,911 It is just looking at your own documents that you care about. 81 00:06:15,511 --> 00:06:15,711 Right? 82 00:06:15,951 --> 00:06:38,671 So coming back to this situation where this employee doesn't know what to do, and if an LLM chatbot can basically look at a 200-page document and give him a checklist of things that he needs to do on day one, like dry the seeds in this way, go to this plant, follow this process, follow this conditioning process, follow this quality process, those kind of things will help him get started. 83 00:06:42,191 --> 00:06:44,591 Now, what are the gaps that we have tried to address? 84 00:06:44,631 --> 00:06:47,151 I mean, talk about the traditional LLM gaps. 85 00:06:47,151 --> 00:06:51,551 It's a large language model, like a ChatGPT model that you've used, cloud model that you've used. 86 00:06:51,711 --> 00:06:53,871 It does not know anything about your enterprise data. 87 00:06:54,111 --> 00:06:56,831 It knows a lot about the rest of the world's data. 88 00:06:57,271 --> 00:07:03,711 When you ask questions to the LLM in general, it can give you generic responses, not tied to your own process. 89 00:07:04,191 --> 00:07:12,031 But with the RAG in the enterprise, that fixes that problem because then now you have your LLM pointing to your data and not the rest of the world. 90 00:07:12,591 --> 00:07:20,751 And with Azure AI Foundry, the contract with Microsoft is such that your data will not leave or the LLM will not be trained on your data, right? 91 00:07:20,911 --> 00:07:24,031 Or will not be retrained on your data outside your enterprise. 92 00:07:25,151 --> 00:07:26,471 So we solved that problem. 93 00:07:26,471 --> 00:07:30,511 There is 80% less hallucination because it is looking at a specific context. 94 00:07:31,071 --> 00:07:35,231 10 to 100 times more cheaper because you don't have to retrain the model each time. 95 00:07:36,031 --> 00:07:38,431 It takes 2 years to retrain a large language model. 96 00:07:38,671 --> 00:07:42,271 But when you give it enough context, it's just going to give you a meaningful response. 97 00:07:42,591 --> 00:07:47,391 And you can also fine tune your LLM in Azure AI Foundry, which I'll show you in a bit. 98 00:07:48,671 --> 00:07:50,431 Very, very cheap, in a very cheap way. 99 00:07:50,591 --> 00:07:52,351 You don't have to spend millions of dollars for that. 100 00:07:53,711 --> 00:07:55,151 Day's time to do production. 101 00:07:55,151 --> 00:08:08,511 Basically, you can go to production in a matter of, so at Cortiva, if someone comes up with a new use case, like a training, onboarding, knowledge assistance, we can onboard them on their data with a rank-based profile in four hours. 102 00:08:10,111 --> 00:08:12,271 Full audit trail of who's asking what questions. 103 00:08:12,351 --> 00:08:14,191 There are safeguards and guardrails around it. 104 00:08:14,191 --> 00:08:15,631 There are evaluations that we do. 105 00:08:15,871 --> 00:08:19,551 If someone is asking a harmful question, we flag that, right? 106 00:08:19,711 --> 00:08:22,831 Someone is asking an explicit question, we flag them. 107 00:08:23,471 --> 00:08:26,511 So we have all the full audit trails with the Azure AI Foundry. 108 00:08:28,431 --> 00:08:31,471 Now, this is the architecture that I want to go through real quick. 109 00:08:32,351 --> 00:08:33,631 You have your source documents. 110 00:08:34,031 --> 00:08:35,951 They could be living anywhere in your enterprise. 111 00:08:36,191 --> 00:08:39,751 It could be a SharePoint, but the major part of it, it's all going to be on SharePoint. 112 00:08:40,591 --> 00:08:45,231 I believe everyone knows about SharePoint or a knowledge repository somewhere, right? 113 00:08:46,031 --> 00:08:48,831 OneDrive, your physical folders, BLOB containers. 114 00:08:49,231 --> 00:08:53,471 I've seen a lot of documents or unstructured or semi-structured documents living on a BLOB container. 115 00:08:54,591 --> 00:09:02,351 So the way the pipeline works is you extract or vectorize the data that you have. 116 00:09:02,671 --> 00:09:05,791 It could be hundreds of, you know, PDFs or thousands of PDFs. 117 00:09:06,031 --> 00:09:07,711 You vectorize it. 118 00:09:08,031 --> 00:09:14,831 Foundry will automatically chunk it into several chunks with the chunk IDs and then you embed it. 119 00:09:15,511 --> 00:09:17,231 using a large language embedding model. 120 00:09:18,111 --> 00:09:20,991 Then you vectorize and index it through Azure AI search service. 121 00:09:21,071 --> 00:09:22,311 And I'm going to take you through that. 122 00:09:22,311 --> 00:09:24,511 These are just the terms that you just have to remember for now. 123 00:09:24,671 --> 00:09:26,231 I'm going to take you and show you what I mean. 124 00:09:27,871 --> 00:09:40,031 The retrieve, generate, and grounded answers are the last and final pieces, where essentially your question also gets converted to basically a vector. 125 00:09:40,031 --> 00:09:44,031 Who knows, I mean, do you know what a vector index is here? 126 00:09:46,991 --> 00:09:48,431 Any idea what a vector index is? 127 00:09:54,181 --> 00:09:55,941 Okay, so I can explain. 128 00:09:58,021 --> 00:10:04,341 So a vector index is basically, let's assume that you have a set of floating point numbers. 129 00:10:04,341 --> 00:10:08,941 Okay, so if you have, if you have a, if you ask a question, what is the weather outside look like today? 130 00:10:09,631 --> 00:10:19,391 That question gets converted into floating point numbers or coordinates like latitudes and longitudes, which does not make sense to a human, but makes a lot of sense to the computer. 131 00:10:20,031 --> 00:10:24,631 Now your grounding information, which is your source data, is also vectorized. 132 00:10:25,231 --> 00:10:36,911 So your documents that you have uploaded, your enterprise documents, your SOPs, plans, procedures, information manuals, your codes, everything is vectorized in the vector database. 133 00:10:37,391 --> 00:10:40,351 Those vectorizations are nothing but floating point numbers. 134 00:10:40,831 --> 00:10:46,431 So when you ask a question to the chat bot, where can I find a plant operating manual? 135 00:10:46,431 --> 00:10:50,511 Or what is the process of drying a seed? 136 00:10:51,151 --> 00:10:55,631 It basically gets converted into a numerical representation. 137 00:10:56,191 --> 00:11:11,551 and then it looks at, so we have a big floating point number, and it goes back and looks at your vector index of your documents and tries to do a similarity search, which means it tries to go to the number which is closest to your question. 138 00:11:12,031 --> 00:11:14,191 Everything gets converted into a floating point number. 139 00:11:14,671 --> 00:11:19,511 So if you have a number like 999, it's going to look at a range of the closest numbers to 999. 140 00:11:20,031 --> 00:11:22,591 And then it gives you more information and context. 141 00:11:22,951 --> 00:11:31,471 What that means is that you don't, it's basically not doing a keyword search, but it's going to go and search for the meaning inside the document, right? 142 00:11:31,511 --> 00:11:41,871 And then when we combine the hybrid retrieval, that's known as the hybrid retrieval, where you combine the keyword search and the vector search so that it gives you more context and the meaning. 143 00:11:41,951 --> 00:11:51,951 And then we also have another concept of re-ranking on the Foundry, which gives you even better quality of the search outcome because it's a combination of both hybrid plus re-ranking. 144 00:11:59,471 --> 00:12:03,951 So semantic or vector search is incredibly powerful for natural language questions. 145 00:12:04,591 --> 00:12:06,751 It can paraphrase queries and conceptual lookups. 146 00:12:07,311 --> 00:12:13,471 This is what makes the chatbot feels like it actually understands you rather than doing a control F, like finding your information. 147 00:12:14,351 --> 00:12:24,711 And the BM25, which is another concept, is a keyword-based search where it's like a keyword specialist that is looking at blind spot information which are more specific, like a product code. 148 00:12:24,791 --> 00:12:27,551 A vector search will not be good for specific information. 149 00:12:27,871 --> 00:12:32,191 It'll be only good for a contextual-based search or meaningful-based search. 150 00:12:32,431 --> 00:12:34,431 That's where the keyword search comes into picture as well. 151 00:12:34,431 --> 00:12:37,391 And when you combine both these, it becomes even more powerful. 152 00:12:40,351 --> 00:12:42,951 So this is an example of a document intelligence pipeline. 153 00:12:42,951 --> 00:12:50,431 As I said, the source documents, we have something known as Azure Document Intelligence, which basically looks at all your documents, thousands of documents. 154 00:12:50,591 --> 00:12:52,591 There's no limit to it how much documents you can vectorize. 155 00:12:53,471 --> 00:12:57,871 you can then it automatically decides what documents need to be chunked with what tokens. 156 00:12:58,351 --> 00:13:02,751 Now, tokens are number of words that it will basically combine into a chunk. 157 00:13:03,551 --> 00:13:07,791 So, depending on if you have short policies, the token size is chosen accordingly. 158 00:13:08,191 --> 00:13:16,631 By default, I think the Foundry interface picks the default token of 512 tokens, which is approximately 500 words. 159 00:13:16,831 --> 00:13:20,751 It does some metadata enrichment, the security, timestamp, ACL, things like that. 160 00:13:20,791 --> 00:13:27,471 And then the OpenAI embeddings convert that into a vector index and store it in the search service, Azure AI search service. 161 00:13:27,471 --> 00:13:28,671 There are several components. 162 00:13:28,911 --> 00:13:30,911 It's a whole enterprise architecture framework. 163 00:13:31,231 --> 00:13:39,471 Within the Foundry, you are also connecting to an Azure AI search service, which actually does the vectorization process for you. 164 00:13:40,431 --> 00:13:45,071 Foundry hosts all the models and LLMs, which I'll show you in a little bit. 165 00:13:45,991 --> 00:13:47,791 I just wanted the concepts to be clear before. 166 00:13:49,711 --> 00:13:52,991 Now, what is the benefit of using an Azure AI Foundry? 167 00:13:53,071 --> 00:13:54,271 There's a model catalog. 168 00:13:54,351 --> 00:13:55,951 There are different kinds of playgrounds. 169 00:13:55,951 --> 00:13:57,071 There's an agent playground. 170 00:13:57,311 --> 00:13:58,831 There are tools that you can use. 171 00:13:59,871 --> 00:14:01,151 And there is also a chat playground. 172 00:14:01,391 --> 00:14:03,711 There are hubs and projects found. 173 00:14:03,791 --> 00:14:05,071 I'm not going to go into the details. 174 00:14:05,071 --> 00:14:06,751 These are all high-level information. 175 00:14:06,751 --> 00:14:08,471 There is a concept of MCP connectors. 176 00:14:08,471 --> 00:14:12,751 Now, for example, how many of you here know about lake houses? 177 00:14:12,751 --> 00:14:14,991 Do you have lake houses in your 178 00:14:16,071 --> 00:14:16,471 enterprise. 179 00:14:16,471 --> 00:14:25,711 So a lake house is a concept where you bring in data for analytics and machine learning and AI purposes into a central location. 180 00:14:25,711 --> 00:14:31,231 So all your transactional systems, you bring in the data into that main reporting engine. 181 00:14:31,631 --> 00:14:40,431 They used to call it the warehouses, but now the term has changed to a lake house, which is, you know, you can get billions and billions of records in your lake house. 182 00:14:40,591 --> 00:14:43,551 Those are mostly structured data, not unstructured data. 183 00:14:44,111 --> 00:14:49,151 So with MCP connectors in Foundry, you can tap into those like a Fabric lake house. 184 00:14:50,151 --> 00:14:51,951 Have you heard about Fabric? 185 00:14:52,111 --> 00:14:53,111 It's an Azure concept. 186 00:14:53,151 --> 00:14:54,751 Have you heard about Databricks? 187 00:14:54,751 --> 00:14:55,631 Who has heard about Databricks? 188 00:14:57,151 --> 00:15:00,991 Databricks has something known as Genie spaces, which basically connect to structured data. 189 00:15:00,991 --> 00:15:02,751 So you can bring that in data. 190 00:15:02,751 --> 00:15:11,351 So you can marry the structured data, which is the relational data and unstructured data and semi-structured data using AI Foundry. 191 00:15:11,871 --> 00:15:12,751 So it's a platform. 192 00:15:13,671 --> 00:15:15,711 and a tool to build on top of. 193 00:15:18,591 --> 00:15:31,311 So again, this is central IT now, this is the best approach where a central IT team manages the infrastructure for Foundry and enables the different teams to use the Foundry infrastructure. 194 00:15:31,711 --> 00:15:40,511 So it's like provisioning the infrastructure first, and then your project teams can pick and choose the projects they want to create, agents they want to create, right, resources they want to create. 195 00:15:44,031 --> 00:15:57,551 So what powers your rack system is the AI hub, AI projects that we have in the Foundry, the FCP connectors, playgrounds, vector stores, and then of course you have the observability through app insights and open telemetry. 196 00:16:00,511 --> 00:16:04,431 Now this is what how you are going to build a rack-based chatbot. 197 00:16:04,431 --> 00:16:05,231 I'll show you in a demo. 198 00:16:05,551 --> 00:16:08,751 We create the resources for a project, you select a model, 199 00:16:09,151 --> 00:16:10,391 you connect your knowledge. 200 00:16:10,391 --> 00:16:12,351 This is all through zero code. 201 00:16:12,471 --> 00:16:14,191 Okay, you can do everything through Foundry. 202 00:16:14,671 --> 00:16:19,871 So basically, you create a project for Foundry, then you select the model from a catalog of 11,000 models. 203 00:16:20,191 --> 00:16:25,391 When I say model, it is your open source GPT models that we have, LLMs, right? 204 00:16:25,591 --> 00:16:31,551 Like ChatGPT or cloud or deep, you know, Gemini, right? 205 00:16:31,551 --> 00:16:32,991 Meta models, deepseek. 206 00:16:33,391 --> 00:16:34,751 Then you connect your knowledge. 207 00:16:34,751 --> 00:16:38,031 It can be a SharePoint site or it can be a BLOB storage. 208 00:16:38,591 --> 00:16:40,191 You configure the index. 209 00:16:40,191 --> 00:16:42,751 You vectorize and configure your index through Azure AI service. 210 00:16:43,391 --> 00:16:45,311 You can tune and evaluate your models. 211 00:16:45,551 --> 00:16:49,311 You can fine tune your models with your data by giving it a training data set. 212 00:16:49,791 --> 00:16:52,511 And then publish and integrate, which means you can publish your agents. 213 00:16:53,231 --> 00:16:56,431 You would need a custom integration either through a Teams channel. 214 00:16:56,511 --> 00:16:57,391 Have you used Teams? 215 00:16:58,111 --> 00:16:58,431 Yes. 216 00:16:58,591 --> 00:17:01,311 So you can create a Teams channel directly through Foundry. 217 00:17:01,871 --> 00:17:03,391 You don't have to write a single line of code for that. 218 00:17:05,071 --> 00:17:15,071 And then you can also call the APIs that you have created on the Foundry for custom interface, just like a regular API call with a key-based authentication. 219 00:17:16,831 --> 00:17:23,551 So again, this is a production-grade implementation that I've done at Cortiva, and this is just an example of what we have done. 220 00:17:24,271 --> 00:17:29,551 So we use the SSO inbuilt SSO. 221 00:17:30,991 --> 00:17:35,591 sort of for a single sign-on experience and multi-factor authentication from user perspective when he signs in. 222 00:17:35,591 --> 00:17:41,231 On the back end, we have Cosmos DB, which is another, you can store your chat responses and history in any database. 223 00:17:41,231 --> 00:17:43,951 We chose Cosmos DB because it's a big database. 224 00:17:43,951 --> 00:17:46,511 It's, you know, it can store, you know, terabytes of data. 225 00:17:46,911 --> 00:17:57,631 So whatever interactions you have with your chat bot, we wanted to capture all the interactions that user had, what kind of questions they were asking, what kind of ratings they were giving to the responses. 226 00:17:58,071 --> 00:18:00,831 Whether it was a thumbs up or a thumbs down, we wanted to capture all that. 227 00:18:01,871 --> 00:18:02,991 So we use Cosmos DB. 228 00:18:02,991 --> 00:18:07,391 The knowledge part of it, the other part that you see in purple here, it's all based on Foundry. 229 00:18:07,391 --> 00:18:09,071 It's all a part of the Foundry framework. 230 00:18:12,111 --> 00:18:18,271 Now, this is just at a glance that one of the implementations I made at Corteva was a Sprout. 231 00:18:18,311 --> 00:18:19,471 We call it the Sprout. 232 00:18:19,551 --> 00:18:24,591 It's an enterprise-grade chatbot meant for 22,000 global users for our seeds business. 233 00:18:25,391 --> 00:18:27,551 So we have a seeds business and a crop protection business. 234 00:18:27,871 --> 00:18:28,991 This is for the seeds business. 235 00:18:29,471 --> 00:18:35,071 What we found is that there were a couple of use cases that we were able to enable on Foundry. 236 00:18:35,071 --> 00:18:40,591 One of them was an employee who was 25 years with the company, decided to leave the company. 237 00:18:41,151 --> 00:18:44,271 And he had a bunch of documents that he had created. 238 00:18:44,751 --> 00:18:47,951 Those were all semi-structured or unstructured documents. 239 00:18:48,671 --> 00:18:52,351 Now we had to train a new employee in his position in 15 days. 240 00:18:53,151 --> 00:18:55,151 They asked us to create a profile on Sprout. 241 00:18:55,871 --> 00:18:57,471 That's how the whole evolution started. 242 00:18:57,471 --> 00:19:02,191 We were able to create that profile and bring the new person coming in on the chatbot. 243 00:19:02,831 --> 00:19:08,831 That was our first validation that yes, chatbots can do a lot than we think about otherwise. 244 00:19:09,151 --> 00:19:13,551 So he got trained in 15 days and he is right now performing really well. 245 00:19:13,551 --> 00:19:17,991 So the feedback that we are getting that yes, the chatbot was giving accurate answer. 246 00:19:18,511 --> 00:19:22,511 We also were increasing the confidence of the end user by citing each and every source. 247 00:19:22,591 --> 00:19:32,271 So citations are very important when you implement chatbots in your companies that you cite the responses and trace it back so that the users can trace it back to the original document. 248 00:19:32,591 --> 00:19:37,871 That way they have high trust in what they are seeing and it also is a good indicator of their hallucinations. 249 00:19:39,551 --> 00:19:46,431 So we implemented that and that's quickly we created a framework so that we can scale to multiple different use cases. 250 00:19:48,191 --> 00:19:50,031 onboarding assistance. 251 00:19:50,111 --> 00:19:54,671 We have 100, we have at least 10 different kinds of profiles now serving hundreds of customers. 252 00:19:56,271 --> 00:20:01,951 So another profile which was very interesting was there were leaders who had to present something every Monday morning. 253 00:20:02,511 --> 00:20:15,311 And they had these Excel sheets that they couldn't sometimes if they spent hours making diagrams and charts, you know, things like that, you do on Excel, you create a bar chart, you create a pie chart, you create a histogram, all those regular sales report, marketing reports. 254 00:20:16,031 --> 00:20:24,951 So we enabled that profile where they upload the Excel on the chatbot and the chatbot info the metrics that they want and create the charts and graphs for them. 255 00:20:25,311 --> 00:20:31,391 So you could also use it for scenarios other than the regular chatbot experience for more of an agentic AI. 256 00:20:31,871 --> 00:20:35,631 So then we also had, so we used tools like Code Interpreter to enable those things. 257 00:20:36,671 --> 00:20:44,031 We also have other kind of agentic profiles where we are asking them to go to Databricks Genie Spaces and 258 00:20:44,511 --> 00:20:49,631 give back structured data or structured information, like a SQL query, natural language. 259 00:20:49,951 --> 00:20:52,351 We ask questions in natural language, it comes back with a SQL query. 260 00:20:53,231 --> 00:20:58,751 So these are some of the use cases we have enabled at Gotiva. 261 00:20:58,751 --> 00:21:03,631 95%, we've seen a 95% retrieval accuracy because we log and monitor all of our responses. 262 00:21:04,671 --> 00:21:13,151 Less than two second response time and 40 to 60% cost reduction, especially for places like Brazil, 263 00:21:13,631 --> 00:21:23,071 South America or Europe, where the documents that we have are in English, they don't need to translate because AI is a multilingual, LLMs are multilingual. 264 00:21:23,071 --> 00:21:27,311 So they respond back in your own language, not the English language. 265 00:21:27,471 --> 00:21:31,311 So that was very helpful because users are now interacting in their own language there. 266 00:21:31,711 --> 00:21:33,071 In the past, they had to translate. 267 00:21:33,231 --> 00:21:34,591 So we saved a lot of cost there. 268 00:21:36,191 --> 00:21:42,671 Now, what are the situation like, what are the RAGs can be implemented not just in the seed business, but also across the board. 269 00:21:42,911 --> 00:21:47,951 And we have seen if you have gone to several, websites like Amazon uses RAG all the time. 270 00:21:49,071 --> 00:21:58,431 There are all these, you know, chat bots that you see on your credit card, your banks, right, bank apps, they use RAG all the time because they are looking at FAQs and responding back to you. 271 00:21:59,231 --> 00:21:59,551 So, 272 00:22:00,911 --> 00:22:06,911 health care policy about retrieval, manufacturing, maintenance, SOP assistance, troubleshooting guides for your workers. 273 00:22:07,791 --> 00:22:22,511 If you are, if you have a shop floor and you're in the manufacturing business, you have so many processes that the person needs, employee needs to follow to create a product, right, or to basically build a product in your shop floor. 274 00:22:22,751 --> 00:22:24,671 Those things can be automated through Rack-based chatbots. 275 00:22:25,711 --> 00:22:29,311 Logistics, route planning, customs documentation, things like that. 276 00:22:29,551 --> 00:22:31,431 Now, we have one of the profiles where... 277 00:22:31,551 --> 00:22:34,751 And I'll also show you in a demo where it can basically write up something for you. 278 00:22:34,911 --> 00:22:43,391 If you give it enough context that, hey, you know, I have this requirement, can you write up a draft e-mail for me based on the context I'm giving you? 279 00:22:43,391 --> 00:22:46,351 Financial services use compliance bots all the time. 280 00:22:47,151 --> 00:22:48,751 Public sector industries are using it. 281 00:22:49,071 --> 00:22:53,311 Citizen service assistance, guidance, grounded regulations, et cetera. 282 00:22:56,031 --> 00:23:06,111 So benchmarks that we have seen, as I said, 95% retrieval efficiency, 80% hallucination reduction, 40 to 60% cost cutting, and, you know, versus fine-tuning or retraining agents. 283 00:23:07,311 --> 00:23:09,631 These are the some high-level numbers, industry-wide. 284 00:23:11,231 --> 00:23:13,391 Scalability and security at enterprise. 285 00:23:14,111 --> 00:23:19,391 We have, as I said, we can support 10,000 plus concurrent users, less than 100 millisecond vector search. 286 00:23:19,711 --> 00:23:24,351 Bring your own key encryption and 0 public exposure, because it's all self-contained within the enterprise. 287 00:23:26,271 --> 00:23:30,751 Best practices are you have to, so it is a garbage in and garbage out. 288 00:23:30,751 --> 00:23:32,511 Your chatbot is as good as your data. 289 00:23:32,911 --> 00:23:35,871 If your data is crappy, the chatbot is also going to be crappy. 290 00:23:36,351 --> 00:23:39,031 So that's pretty, it's like a no-brainer. 291 00:23:39,231 --> 00:23:50,911 So you have to understand that the data has to be accurate for your chatbots to work accurately, especially in the RAG world where we are telling the LLM that don't use your intelligence, use our intelligence. 292 00:23:51,311 --> 00:23:54,831 Don't retrieve information from your memory, use our data. 293 00:23:55,791 --> 00:23:59,391 So retrieval design, hybrid search, re-ranking are the best practices. 294 00:23:59,711 --> 00:24:04,271 Using a large language model, text embedding large is the best practice currently. 295 00:24:04,671 --> 00:24:08,511 Observability, you have to track what you know, what people are asking questions on. 296 00:24:09,391 --> 00:24:10,551 Security is very important. 297 00:24:14,751 --> 00:24:20,431 We have basically the custom chatbot interface to only people. 298 00:24:20,711 --> 00:24:21,151 the enterprise. 299 00:24:21,231 --> 00:24:23,391 So that is also very important that you don't want exposure. 300 00:24:23,471 --> 00:24:28,351 Rotating those keys are also very important because these are all key-based authentication from your endpoints perspective. 301 00:24:28,511 --> 00:24:31,311 Cost optimization, we have a way to monitor cost. 302 00:24:31,311 --> 00:24:33,711 Now how do you know what's your input and output token cost? 303 00:24:34,191 --> 00:24:40,671 And based on how many number of users are hitting your chatbot every day, there is a visibility for that too on the Foundry interface, which I'll show you. 304 00:24:40,991 --> 00:24:41,711 Automation. 305 00:24:42,351 --> 00:24:43,271 It's very important now. 306 00:24:43,271 --> 00:24:49,911 This might not be important for business users, but from an IT or technical users, you should be automating your deployments. 307 00:24:49,911 --> 00:24:59,231 You should not be creating Foundry infrastructure manually every time, like staging up a search service or staging up a model, creating a model, adding a model. 308 00:24:59,551 --> 00:25:03,951 Those can all be automated through something as a DevOps, IAC, infrastructure as a code. 309 00:25:03,951 --> 00:25:05,871 Have you heard about infrastructure as a code? 310 00:25:07,711 --> 00:25:07,951 Yes. 311 00:25:08,191 --> 00:25:10,511 So these can all be automated for you. 312 00:25:14,111 --> 00:25:16,831 so this is just a slide on trends in Azure. 313 00:25:16,831 --> 00:25:23,391 So what is the current trends are Foundry IQ is very important because it connects to several different kinds of systems. 314 00:25:24,031 --> 00:25:30,911 Foundry IQ can connect to structured data, unstructured data, semi-structured data, laying anywhere in your enterprise. 315 00:25:30,911 --> 00:25:34,351 It just doesn't have to be Microsoft enabled. 316 00:25:34,351 --> 00:25:36,111 It can be also auto databases. 317 00:25:36,111 --> 00:25:38,511 It can be other types of databases. 318 00:25:38,991 --> 00:25:43,231 or other types of disparate systems that you can connect to through MCP servers. 319 00:25:44,671 --> 00:25:59,711 Agent AI is very, basically Microsoft is now focused on doing agentic rack-based systems where you have an agent which basically calls another agent and then they basically synchronize with each other and come back with the information. 320 00:25:59,871 --> 00:26:05,471 That's where they're going to, the multi-supervisory agent that can delegate the tasks and come back with the accurate response. 321 00:26:06,991 --> 00:26:13,311 I'll show you an example of that as well, how we can do that easily through a workflow in Foundry. 322 00:26:14,151 --> 00:26:17,791 Efficient models, again, this is just for information purposes. 323 00:26:17,791 --> 00:26:21,471 There are some smaller, large language models that you can use for smaller tasks. 324 00:26:21,631 --> 00:26:23,871 You don't have to use a large language model all the time. 325 00:26:24,031 --> 00:26:26,271 There's also a small LLM that you can use. 326 00:26:29,191 --> 00:26:35,151 Like streaming or even indexes can be updated at a time or at a frequency. 327 00:26:39,551 --> 00:26:43,471 Now, go to the most important section, which is the demo, which is what I think you guys are waiting for. 328 00:26:43,951 --> 00:26:46,351 So, let me show you how a rack. 329 00:26:46,351 --> 00:26:50,031 So, first starting with what's a vector, how does a vector index looks like? 330 00:26:50,031 --> 00:26:53,311 So, if you see my screen here, this is how a vector index looks like. 331 00:26:53,311 --> 00:27:02,191 This is a IMDB movie data set, and this is a this is a set of 5000 best or top movies later on IMDB. 332 00:27:03,191 --> 00:27:04,751 And if you see, these 333 00:27:05,231 --> 00:27:07,151 are the vectors for that. 334 00:27:07,231 --> 00:27:10,431 And if you see, these are the chunk IDs, chunk title. 335 00:27:10,431 --> 00:27:12,191 So it's a CSV file. 336 00:27:12,191 --> 00:27:14,591 Out of that CSV file, we have so many vectors. 337 00:27:14,911 --> 00:27:18,031 So I was telling you that vector is a floating point numerical value. 338 00:27:18,831 --> 00:27:24,351 And every word that's in that CSV file gets converted into a vector or a numerical floating point. 339 00:27:24,991 --> 00:27:29,791 And when you ask a question that, hey, give me the list of top 10 highly rated movies, 340 00:27:30,031 --> 00:27:35,631 it's going to go and convert that to a numerical representation and find the closest match in this vector database. 341 00:27:37,711 --> 00:27:42,271 This is the IMDB movie data set that's hosted on a BLOB container. 342 00:27:43,311 --> 00:27:50,031 And I'll show you, we have just vectorized that here as a knowledge source. 343 00:27:50,791 --> 00:27:53,071 So, this is the index I was talking about. 344 00:27:53,071 --> 00:27:58,751 The first step is to do the grounding data, which is the grounding data is actually your enterprise documents. 345 00:27:58,831 --> 00:28:07,551 Now, you can automate this by creating some pipelines where you can extract data from your SharePoint at a frequency, because there could be hundreds of new documents every day. 346 00:28:07,951 --> 00:28:09,471 You don't have to do this manually. 347 00:28:09,711 --> 00:28:12,991 I'm showing it for the demo purposes, but in enterprise scenarios. 348 00:28:13,311 --> 00:28:20,831 You should create a document ingestion pipeline if you have a lot of unstructured data and then automate the vector indexing part of it. 349 00:28:22,271 --> 00:28:24,431 So the way I created this index was pretty simple. 350 00:28:24,831 --> 00:28:26,831 I imported the data from BLOB container. 351 00:28:28,351 --> 00:28:30,471 Yeah, I'm out of the quota. 352 00:28:30,591 --> 00:28:31,431 So I have 3 indexes. 353 00:28:31,431 --> 00:28:32,831 That's why it doesn't allow me, but that's OK. 354 00:28:33,471 --> 00:28:40,191 So the way I do it is basically I've created these indexes and this tier of search service, which is a basic tier, 355 00:28:40,591 --> 00:28:42,591 has only three indexes that I can create. 356 00:28:42,591 --> 00:28:52,271 So I couldn't show it to you, but there are these three indexes that I've created through an import process, which is a manual process, and you can automate that easily with very little code. 357 00:28:52,911 --> 00:28:59,151 So now what I'm going to do is I'm going to create an agent where we type agent to this particular IMDB index. 358 00:29:00,511 --> 00:29:01,311 Right, let's call it. 359 00:29:01,391 --> 00:29:02,431 So this is an OK. 360 00:29:02,431 --> 00:29:03,791 This is the Foundry interface. 361 00:29:04,111 --> 00:29:10,111 Now coming back to Foundry, the I think I showed you the Azure AI search service. 362 00:29:10,671 --> 00:29:11,791 This is what I was talking about. 363 00:29:12,191 --> 00:29:14,991 And then the BLOB container and the vector store. 364 00:29:15,231 --> 00:29:26,751 We go to the foundry interface, which is where the playground is for creating agents and fine-tuning models and creating evaluations and looking at the tools and things like that. 365 00:29:27,151 --> 00:29:28,991 So if you see, there are several tools available. 366 00:29:28,991 --> 00:29:31,871 Let me see, let me discover. 367 00:29:34,751 --> 00:29:36,991 These are the models that you can choose and pick and choose from. 368 00:29:38,551 --> 00:29:40,271 Let's show you some of these models here. 369 00:29:40,671 --> 00:29:45,991 There are more than 11,000 models, and every day there are at least 2 or 3,000 models that are being added. 370 00:29:45,991 --> 00:29:51,631 Here we go. 371 00:29:52,351 --> 00:29:56,031 So cloud is there, and then you have GPT 5.4. 372 00:29:56,031 --> 00:29:57,471 So these are all available to you. 373 00:29:57,671 --> 00:30:05,311 And then the data, as I said, these models are available exclusively for your subscription or tenant in your organization. 374 00:30:05,391 --> 00:30:06,591 They are not going to train 375 00:30:06,991 --> 00:30:10,751 these models with your data and expose it to the outside world. 376 00:30:10,831 --> 00:30:12,911 That's why we have Foundry, right? 377 00:30:12,911 --> 00:30:14,591 Because everything is self-contained in the framework. 378 00:30:17,071 --> 00:30:19,631 So let's go back and build an agent. 379 00:30:21,631 --> 00:30:24,031 Let's create an IMDB agent, movie advisor agent. 380 00:30:31,601 --> 00:30:32,721 Now what's the first step? 381 00:30:32,721 --> 00:30:36,561 The first step is to connect your data source, right? 382 00:30:36,881 --> 00:30:38,561 Because how does the agent know what to do? 383 00:30:39,231 --> 00:30:43,311 If you don't give it the data, the grounding data, so there is this knowledge section here right here. 384 00:30:47,071 --> 00:30:48,991 So it's the Foundry IQ I was talking about. 385 00:30:48,991 --> 00:30:50,431 You can connect any kind of services. 386 00:30:50,431 --> 00:30:53,551 I want to connect my search service to the knowledge base. 387 00:30:53,551 --> 00:30:56,591 If you see, I have my index that I was showing you listed here. 388 00:30:56,911 --> 00:31:00,831 I connect my index to this Foundry IQ. 389 00:31:00,911 --> 00:31:06,431 I automatically have a web search here because I also want to go and look at rotten tomatoes. 390 00:31:06,831 --> 00:31:11,711 And whatever movie recommendations I get from IMDB data set, I want to validate that with Rotten Tomatoes. 391 00:31:11,871 --> 00:31:13,311 You know Rotten Tomatoes, right? 392 00:31:13,791 --> 00:31:16,591 It's just a public crowdsourcing for movies. 393 00:31:16,911 --> 00:31:20,991 So now I have instructions here that I have predefined. 394 00:31:20,991 --> 00:31:22,351 I don't want to create them by hand. 395 00:31:22,671 --> 00:31:27,871 So now these instructions are very important because this tells the agent on what to do. 396 00:31:28,351 --> 00:31:32,671 So if you read this in a nutshell, what it tells you is 397 00:31:32,991 --> 00:31:40,711 You are an intelligent movie discovery assistant powered by a vectorized knowledge base of 5,000 movies from the TMDB. 398 00:31:40,711 --> 00:31:46,911 Now, the data set, they call it the movies database because of copyright infringement, right? 399 00:31:47,471 --> 00:31:52,271 So the data set I have calls it TMDB, but it's actually IMDB data. 400 00:31:52,671 --> 00:31:59,871 So it, and then we tell it to basically go and 1st look at our knowledge source and then go to 401 00:32:00,391 --> 00:32:07,871 do a live web search to Rotten Tomatoes for real-time scores, audience scores, critics, et cetera, movie reviews. 402 00:32:09,151 --> 00:32:10,031 These are the instructions. 403 00:32:10,031 --> 00:32:10,911 These are very important. 404 00:32:11,231 --> 00:32:14,671 The model will exactly, and this is an exercise on its own. 405 00:32:14,671 --> 00:32:27,151 When we created these profiles for employees at, I mean, for our user base at Corteva, we had to go through and talk to domain experts for each and every profile that we were creating that, hey, how do you want this profile to behave? 406 00:32:27,471 --> 00:32:29,471 Every profile is going to behave in a different way. 407 00:32:30,111 --> 00:32:32,231 And it's all configured through this prompt template. 408 00:32:32,231 --> 00:32:35,151 This is a very important, this is known as a system prompt engineering. 409 00:32:35,471 --> 00:32:44,351 This is a very important aspect in the way that you tell the agent in natural language of how it needs to behave or what persona it needs to take. 410 00:32:45,631 --> 00:32:47,191 When someone asks a specific question. 411 00:32:47,191 --> 00:32:52,311 So we also have examples here and the outputs, how it should generate the output. 412 00:32:52,591 --> 00:32:57,631 Like when you generate the output, show the title, generate, director, cast, et cetera, et cetera, et cetera. 413 00:32:58,671 --> 00:32:58,911 right? 414 00:32:59,391 --> 00:33:04,671 So special capabilities, semantic discovery, and I give it more specific information, right? 415 00:33:04,991 --> 00:33:09,751 So then there's also voice mode, which I've not, I mean, it's cool to play with, but I've not played around with it too. 416 00:33:09,751 --> 00:33:17,151 This is a new feature where you can basically speak to it and it'll respond back and also speak back to you. 417 00:33:17,391 --> 00:33:19,191 So this is all out-of-the-box, guys. 418 00:33:19,191 --> 00:33:20,271 You don't have to code for it. 419 00:33:20,351 --> 00:33:24,111 You don't have to do like graphs and lang chains and write anything new. 420 00:33:25,151 --> 00:33:31,431 You can, and then the other part which is interesting is you can directly publish the endpoint or you can publish to teams. 421 00:33:31,511 --> 00:33:33,831 and Microsoft 365 as an agent. 422 00:33:34,591 --> 00:33:35,791 So this is out-of-the-box. 423 00:33:36,271 --> 00:33:38,871 And it requires very little technical insight. 424 00:33:38,871 --> 00:33:41,351 We just need to know how to set it up the first time. 425 00:33:41,551 --> 00:33:43,471 That took us quite a while to figure out. 426 00:33:44,111 --> 00:33:47,231 So then there is also memory. 427 00:33:47,231 --> 00:33:53,191 Now some people like to also keep like the past chat history memory of the chat bot based on the 428 00:33:53,631 --> 00:33:57,231 The questions that people have been asking, it goes back and looks at the memory. 429 00:33:57,231 --> 00:34:04,351 That's why if you have seen, if you go to Claude or ChatGPT, the longer the conversations you have with it, the slower it becomes. 430 00:34:04,431 --> 00:34:07,231 Because it has to go back to your past questions and respond back. 431 00:34:08,511 --> 00:34:10,271 Then there is guardrails I was talking about. 432 00:34:10,911 --> 00:34:21,631 Guardrails are nothing but safety features that people, if they are asking self-harming questions, questions that they should not be asking, questions about hacking, unethical stuff, we should block it. 433 00:34:22,671 --> 00:34:24,591 These are all configured by default for you. 434 00:34:26,191 --> 00:34:30,271 So let's save this and let's ask a few questions here real quick. 435 00:34:30,271 --> 00:34:34,151 I think I had a list of questions that I had prepared. 436 00:34:34,151 --> 00:34:34,591 Here we go. 437 00:34:37,551 --> 00:34:40,111 So now the agent is practically ready. 438 00:34:40,111 --> 00:34:42,511 I have pointed it to a GPT 4.1 model. 439 00:34:42,511 --> 00:34:44,111 I can change the model any time I want. 440 00:34:44,111 --> 00:34:44,431 Look at this. 441 00:34:44,831 --> 00:34:46,511 I can change to any model I want. 442 00:34:47,111 --> 00:34:51,871 And then when I ask a question here, let's see what it comes back with. 443 00:34:52,671 --> 00:34:56,831 Hopefully not an error, because I've seen a lot of errors in my demos. 444 00:34:59,071 --> 00:34:59,791 Okay, here we go. 445 00:35:05,991 --> 00:35:10,871 So it gave me all the top five movies in the exact format I had asked him to do. 446 00:35:11,271 --> 00:35:16,311 And also, this tool always approved this tool. 447 00:35:17,831 --> 00:35:20,711 Recommend 5 movies similar to The Matrix, right? 448 00:35:20,791 --> 00:35:23,271 So Matrix was a sci-fi movie. 449 00:35:23,831 --> 00:35:24,711 Inception, 450 00:35:25,791 --> 00:35:26,831 I think these are accurate. 451 00:35:28,351 --> 00:35:31,551 Ghost in the shell, equilibrium, minority report, 13th floor. 452 00:35:31,551 --> 00:35:31,871 Perfect. 453 00:35:31,871 --> 00:35:32,751 So this is bang on. 454 00:35:32,751 --> 00:35:34,111 This is exactly what we wanted. 455 00:35:35,711 --> 00:35:38,111 Let's talk about this. 456 00:35:38,271 --> 00:35:39,391 is my favorite movie. 457 00:35:39,871 --> 00:35:42,271 Why is Shawshank Redemption rated so high? 458 00:35:43,151 --> 00:35:44,191 9.2 or something. 459 00:35:45,151 --> 00:35:45,791 I guess. 460 00:35:46,111 --> 00:35:46,991 So let's see. 461 00:35:50,791 --> 00:35:56,191 Now if you choose a 5.4 model, it performs better in terms of responses, accuracy and all that stuff. 462 00:35:56,231 --> 00:35:59,791 So now it also tells you how much tokens it is used here. 463 00:36:00,031 --> 00:36:03,671 So you know how the, what is an input token and an output token? 464 00:36:03,671 --> 00:36:06,751 Do you guys, are you guys aware of what is an input token and an output token? 465 00:36:07,151 --> 00:36:13,791 Input token is when you ask a question, your, the length of your characters in your question is, you know, gets converted to a token. 466 00:36:14,751 --> 00:36:18,031 So in this case, our input token was 2884 characters. 467 00:36:18,671 --> 00:36:20,431 The output is what it comes back with. 468 00:36:20,591 --> 00:36:22,511 So 343 tokens for the output. 469 00:36:22,511 --> 00:36:25,751 Sometimes it also caches the previous output and shows some response. 470 00:36:25,751 --> 00:36:27,311 So then you save cost that way. 471 00:36:28,591 --> 00:36:31,311 So your cost for using the LLM is monitored this way. 472 00:36:34,431 --> 00:36:37,431 Now, emotional resonance, strong performance, clean writing, et cetera. 473 00:36:37,431 --> 00:36:39,791 So give you some valid responses back there. 474 00:36:41,231 --> 00:36:42,431 Let's do another one. 475 00:36:45,711 --> 00:36:46,591 You see. 476 00:36:50,991 --> 00:36:56,911 Which movie has the highest gross collection, right? 477 00:36:57,951 --> 00:37:02,591 Or made the highest money? 478 00:37:06,511 --> 00:37:10,231 I can ask it, grammatically incorrect questions and it'll still come back. 479 00:37:10,231 --> 00:37:15,951 That's what I want to show you, that even if my grammar is incorrect, it will still come back with the right responses. 480 00:37:20,671 --> 00:37:22,591 Yeah, Avatar, which is accurate. 481 00:37:22,831 --> 00:37:25,151 So 2.9 billion at the box office. 482 00:37:26,351 --> 00:37:35,871 Now, I can tell him that go to Rotten Tomatoes and show me the critic reviews for Avatar. 483 00:37:35,871 --> 00:37:38,511 So that's a public site. 484 00:37:38,671 --> 00:37:45,151 If you asked me to go to a private site, what I did, you were to use a critic. 485 00:37:47,311 --> 00:37:48,671 So that's a web search. 486 00:37:48,871 --> 00:37:51,071 I'm doing a Google search or Bing search, right? 487 00:37:51,711 --> 00:38:01,791 But if I have to query my own internal site for a tool, tool-based search, first of all, you'll have to configure an endpoint, like an API. 488 00:38:02,351 --> 00:38:12,031 And my identity that I will use is a custom, so I'm not going to, so this is a playground for testing, but I will have a custom web app or an interface. 489 00:38:12,351 --> 00:38:16,111 And the app and interface needs to have access to the API. 490 00:38:18,351 --> 00:38:22,991 Yeah, So yeah, it gives me back. 491 00:38:23,231 --> 00:38:27,391 Now I'm going to go back quickly and show you some other agents that I've built. 492 00:38:28,111 --> 00:38:31,071 So this was the agent that we built for scratch in like 10 minutes. 493 00:38:31,631 --> 00:38:35,351 And I have just five more minutes, so I'm just going to do this quick. 494 00:38:35,351 --> 00:38:37,871 I have another agent where it's basically doing something different. 495 00:38:37,871 --> 00:38:40,271 It's basically had a data set of diseases and predictions. 496 00:38:40,751 --> 00:38:45,471 It's a prediction data set where every diseases have certain symptoms. 497 00:38:46,031 --> 00:38:49,631 And we can now predict the disease based on the symptoms that you have. 498 00:38:49,791 --> 00:38:53,631 Or like, for example, if I have a cough and cold, what could I be suffering from, right? 499 00:38:54,031 --> 00:38:56,111 If I have high fever, do I have malaria? 500 00:38:56,191 --> 00:38:57,711 We can ask these kind of questions to the data. 501 00:38:58,551 --> 00:39:04,351 And then there's also a web search that it goes to WebMD to validate that information as another check. 502 00:39:05,391 --> 00:39:07,951 So let's take this prompt, for example. 503 00:39:08,391 --> 00:39:10,991 Which diseases match these symptoms most closely? 504 00:39:10,991 --> 00:39:12,831 Fatigue, cough, and chest pain? 505 00:39:13,311 --> 00:39:14,351 Let's ask this question. 506 00:39:17,191 --> 00:39:22,111 And again, as you see, I have this instructions clearly defined on how the agent needs to behave. 507 00:39:22,111 --> 00:39:25,551 If you don't give it the instructions, it's going to behave very randomly. 508 00:39:28,911 --> 00:39:31,911 It just will behave like a Google search, not exactly the way you want. 509 00:39:31,911 --> 00:39:32,191 Yes, sir. 510 00:39:32,191 --> 00:39:35,151 So for this scenario, what does the data set do? 511 00:39:35,551 --> 00:39:38,031 is it an Excel spreadsheet or is it a markup file? 512 00:39:38,591 --> 00:39:45,791 This is, in this case, it's a CSV file, which has, which has a, or an Excel or a CSV, which has columns and rows. 513 00:39:46,191 --> 00:39:50,831 And every disease is mapped based on symptom severity in the scale of 1 to 10. 514 00:39:50,831 --> 00:39:54,431 You don't have to do anything in between the data set and. 515 00:39:55,831 --> 00:39:56,991 We don't have, that's the beauty. 516 00:39:56,991 --> 00:40:02,591 You don't have to create a very specific machine learning model for this AI to look, it will infer on its own. 517 00:40:02,991 --> 00:40:05,231 That's why the need of machine learning is eliminated. 518 00:40:05,871 --> 00:40:11,231 So I have done a project where we had to estimate costs based on the past spend. 519 00:40:11,951 --> 00:40:19,871 There was an Excel sheet with 10 years worth of data for all the projects that we have spent money on and the scale and the size of the projects. 520 00:40:20,191 --> 00:40:25,791 So without any machine learning technology, I created an agent and gave it that data set. 521 00:40:26,591 --> 00:40:41,151 and I asked and I prompted the agent to behave in a certain way that go and look back all the projects which is similar to my prompt that I'm asking a question on and compare and see what would be my estimated cost in the future if I were to do a similar project. 522 00:40:41,631 --> 00:40:47,471 So it gave me a prediction based on my past data. 523 00:40:49,631 --> 00:40:55,591 Similar to in the past, you had to write a machine learning machine, not with AI, because it has a capability. 524 00:40:57,991 --> 00:40:59,111 Let me ask another question. 525 00:40:59,111 --> 00:41:03,551 So that the with the have some safeguards or guardrails in it, will block certain prompts. 526 00:41:03,871 --> 00:41:10,351 If it has PHI or if it thinks that there is like a health information that he's, you know, I want to retrieve back. 527 00:41:11,551 --> 00:41:12,311 I think this should work. 528 00:41:12,591 --> 00:41:16,031 Which symptoms most strongly categorize diabetes? 529 00:41:16,031 --> 00:41:21,391 So yeah, so based on the data set, we'll first look at the data set and then it'll tell you. 530 00:41:24,591 --> 00:41:31,311 I mean, look at the third question here: which symptoms overlap between typhoid and malaria? 531 00:41:53,181 --> 00:41:54,181 Here we go. 532 00:41:54,181 --> 00:41:58,461 It has, you know, vomiting, high fever, etcetera, etcetera, and then severe context, right? 533 00:41:59,101 --> 00:42:03,101 So, basically, this is a quick demo on the disease prediction agent. 534 00:42:03,101 --> 00:42:05,741 The last and important demo was for the HR. 535 00:42:05,741 --> 00:42:08,861 Now, every enterprise has HR, right? 536 00:42:09,391 --> 00:42:21,391 So now imagine HR gets hundreds and hundreds of resumes every single day, where there is a job opening, they have a database of all the resumes of current employees and newer prospective employees. 537 00:42:22,351 --> 00:42:28,351 Imagine there is a new job requirement that the HR is asked to look for candidates. 538 00:42:29,951 --> 00:42:38,591 Will the HR go and look at each and every resume, download it from a SharePoint site or from an external custom site and review each and every resume? 539 00:42:38,831 --> 00:42:42,751 If she has to review 100 of resumes, she's going to take a week or so to do that. 540 00:42:43,671 --> 00:42:45,631 What if we build an agent that can do that? 541 00:42:46,671 --> 00:42:52,511 Which means the agent will go and vectorize each and every resume that's being added to a specific SharePoint site. 542 00:42:54,031 --> 00:42:58,991 And when you ask a question like, give me the top five data engineering profile, 543 00:43:01,151 --> 00:43:08,031 or the top five data engineering candidates that matter job solution. 544 00:43:08,151 --> 00:43:14,911 And then you can also ask the agent to go to LinkedIn and verify whether that candidate exists or does not exist. 545 00:43:15,311 --> 00:43:21,951 You can also go and do a background check, go to a third-party API to do a background check based on his first, last name and the address. 546 00:43:22,271 --> 00:43:24,631 See if he has a criminal record or things like that. 547 00:43:24,631 --> 00:43:26,671 So all those things can be enabled through Foundry. 548 00:43:27,711 --> 00:43:33,471 So here an example is, find me the qualified data engineers for a senior data breaks role. 549 00:43:33,471 --> 00:43:35,871 For example, this is my requirement, job description or requirement. 550 00:43:40,751 --> 00:43:43,791 Yeah, it's telling, I am the top candidate here. 551 00:43:44,351 --> 00:43:44,911 So that's great. 552 00:43:45,351 --> 00:43:48,751 It's actually, yeah, since I built the agent, so it knows that. 553 00:43:50,591 --> 00:43:53,711 So again, I think this is the last part of the demo. 554 00:43:53,711 --> 00:44:01,631 And then I just want to show how powerful AI can be if you use it the right way with the right safeguards and with the right guard rails. 555 00:44:03,071 --> 00:44:05,871 I hope you enjoyed this and I hope you learned a lot from this. 556 00:44:06,111 --> 00:44:06,511 Thank you. 557 00:44:06,831 --> 00:44:06,991 Yeah. 558 00:44:06,991 --> 00:44:12,591 Do different models have different efficiencies as far as token use and that sort of thing? 559 00:44:13,311 --> 00:44:19,471 Yeah, so the models have a higher efficiency for larger token use compared to a small LM models. 560 00:44:25,151 --> 00:44:30,911 Yes, the newer models are more efficient, like Claude is by far what I've tested for Claude Opus. 561 00:44:31,631 --> 00:44:41,471 It's by far the best model out there for all kinds of tasks, not just chat, but also creating charts and graphs or writing code or doing some extraordinary tasks for you. 562 00:44:43,391 --> 00:44:49,071 But they are more costly, 7.5 times more costly than the low cost model, like GPT ones. 563 00:44:50,711 --> 00:44:51,351 Any other questions? 564 00:44:51,351 --> 00:44:52,191 You have time for Q&A? 565 00:44:52,751 --> 00:44:54,351 Gail, do we have time for Q&A? 566 00:44:56,111 --> 00:44:59,231 OK, yeah, you can talk separately outside this. 567 00:44:59,311 --> 00:44:59,631 It's fine. 568 00:44:59,791 --> 00:45:00,191 Thank you.