1 00:00:00,261 --> 00:00:02,861 Thank you for being here in the production and operations track. 2 00:00:02,861 --> 00:00:03,821 My name is Jake Behrens. 3 00:00:03,821 --> 00:00:07,061 I'll be helping moderate the room here for these sessions today. 4 00:00:07,061 --> 00:00:09,941 And I have the honor of introducing our speakers in here as well. 5 00:00:09,941 --> 00:00:11,181 So good morning. 6 00:00:11,181 --> 00:00:12,581 It's my pleasure to introduce Dr. 7 00:00:12,581 --> 00:00:17,461 Aditya Blue, a data scientist at the Translational AI Center at Iowa State University here. 8 00:00:17,461 --> 00:00:21,501 So with over 15 years of experience in artificial intelligence and machine learning, Dr. 9 00:00:21,501 --> 00:00:23,621 Blue has contributed extensively to the field. 10 00:00:24,181 --> 00:00:32,421 with research published in leading venues such as Nature Computational Science, ICML, and Engineering Applications of Artificial Intelligence. 11 00:00:32,501 --> 00:00:40,181 So his work focuses on applying advanced AI techniques to engineering and manufacturing challenges, bridging cutting-edge research with real-world impact. 12 00:00:40,181 --> 00:00:44,901 So in addition, he plays an active role in developing teaching AI and machine learning programs. 13 00:00:45,301 --> 00:00:48,021 to help translate these innovations into practice. 14 00:00:48,021 --> 00:00:49,701 So in today's session, Dr. 15 00:00:49,701 --> 00:01:02,541 Bhalu will explore the emerging role of tabular foundational models in manufacturing, highlighting how these approaches can simplify workflows, improve predictions, and open new possibilities for industrial AI applications. 16 00:01:02,541 --> 00:01:04,301 So please join me in welcoming Dr. 17 00:01:04,301 --> 00:01:05,061 Aditya Bhalu. 18 00:01:05,701 --> 00:01:05,981 Thanks. 19 00:01:06,101 --> 00:01:06,861 Thanks a lot, Jake. 20 00:01:06,861 --> 00:01:08,741 I hope I'm audible. 21 00:01:10,421 --> 00:01:11,461 So, good morning, everyone. 22 00:01:11,901 --> 00:01:14,901 Thanks, Jake, for the nice introduction. 23 00:01:16,101 --> 00:01:19,781 So, today I'll be talking about tabular foundation models for manufacturing. 24 00:01:21,541 --> 00:01:24,261 Just to give you the context of how this comes in, right? 25 00:01:24,261 --> 00:01:44,501 So I'm sure you all have seen some kind of a tabular data, a bunch of inputs, probably an output, you want to predict whether, what is the manufacturing rate you have or what is the prediction of, when you're going to get your order from Amazon, whether it's going to come tomorrow or day after and all. 26 00:01:44,741 --> 00:01:48,101 So each of them, they're all tabular data in some form or the other. 27 00:01:48,501 --> 00:01:48,981 And 28 00:01:49,621 --> 00:01:58,661 Have you ever wondered, can I just dump all this information into ChatGPT and, ask it to give me the predictions and get done with it? 29 00:01:58,661 --> 00:02:00,821 Has anyone tried something like that so far? 30 00:02:01,701 --> 00:02:02,901 Okay, at least one. 31 00:02:03,941 --> 00:02:10,101 But I'm sure, do you have any experience you want to share in terms of, you know, what you saw when you did that? 32 00:02:10,101 --> 00:02:11,901 Yeah, so I'm actually just a freshman. 33 00:02:12,261 --> 00:02:16,701 I would say some of the more basic ones would be like, maybe I'm working on a problem. 34 00:02:17,461 --> 00:02:18,901 a lot of data or parts to it. 35 00:02:19,101 --> 00:02:23,701 I try to upload it all at once and sometimes it gets lost or it doesn't do the right calculation. 36 00:02:24,221 --> 00:02:30,701 Sometimes I might break it up and sometimes helps, but other times if it's just too much, doesn't work. 37 00:02:30,701 --> 00:02:30,821 Yeah. 38 00:02:30,821 --> 00:02:45,221 So as you mentioned, right, so the too much data is one thing, but fundamentally there's one problem with, you know, dumping all the tabular data to, you know, ChatGPT or any of the any of your favorite LLMs today that you're using. 39 00:02:45,701 --> 00:02:49,221 The problem is that they are not meant for tabular data. 40 00:02:49,621 --> 00:02:50,901 They are meant for language. 41 00:02:50,941 --> 00:02:52,981 They are meant to have a conversation with you. 42 00:02:53,541 --> 00:02:58,821 And they're usually called as large language models or foundation models for language, per se. 43 00:03:00,581 --> 00:03:08,981 But at the same time, you know, in the recent last few years, there's been a lot of shift towards, you know, building foundation models for tabular data. 44 00:03:09,541 --> 00:03:14,181 And that's what is the term tabular foundation models that you're seeing here. 45 00:03:16,661 --> 00:03:35,541 If I want to explain in comparing what is a foundation model to what we do in traditional machine learning models, in traditional machine learning you have a data set, you have a new data set, and essentially you train a fresh model, train them, tune the hyper-parameters, validate, and then deploy. 46 00:03:35,941 --> 00:03:39,141 And you do this for almost every task that you are trying to work on. 47 00:03:39,981 --> 00:03:43,541 Quite often than not, you have to train the models from scratch. 48 00:03:43,541 --> 00:03:45,221 You have to do a lot of feature engineering. 49 00:03:45,701 --> 00:03:55,341 And quite often than not, anyone who has worked with real-world data, they know that the most important thing that you see is cleaning up the data. 50 00:03:55,341 --> 00:03:58,021 There are a lot of cases where you have missing data. 51 00:03:58,341 --> 00:04:08,061 You have cases where the data is probably noisy, and you have to do a lot of data imputation and things like that to even before get started to train the model. 52 00:04:09,341 --> 00:04:16,981 And what we saw as an opportunity is the foundation models, right? 53 00:04:17,301 --> 00:04:25,381 So, think about it today that in ChatGPT, even if you write gibberish, not perfectly grammatical sentences. 54 00:04:25,781 --> 00:04:43,381 and even perhaps with a lot of typos and everything, it still understands what you're saying and it still tries to figure out what you're trying to do and then gives you an output which probably is relevant to unless and until you are giving completely and expect some output, it may not. 55 00:04:43,621 --> 00:04:47,381 But as long as it's reasonably okay, it gives you somewhat reasonable output, right? 56 00:04:47,781 --> 00:04:49,461 So similarly, 57 00:04:49,941 --> 00:05:04,181 Even if you think of the case of Tableau Foundation Models, if there is some noise, if there is some cases where there's some missing data, there's a lot of these issues, you don't have to clean up yourself. 58 00:05:04,661 --> 00:05:07,301 That is 1 relief that you'll see. 59 00:05:07,781 --> 00:05:21,221 And the next part is obviously just like in ChatGPT, you don't for say a medical application, you don't train a separate ChatGPT model or for finance application, you don't train your own model directly. 60 00:05:21,621 --> 00:05:30,981 At least you can fine tune them later on and other things, but most often than not, you know, just using ChatGPT alone gets you quite far along. 61 00:05:32,101 --> 00:05:33,301 It is the same idea. 62 00:05:33,741 --> 00:05:37,061 which you can think of in tabular foundation models as well. 63 00:05:37,061 --> 00:05:47,541 But there is one foundation model that is pre-trained with a lot of synthetic data from different possibilities. 64 00:05:47,541 --> 00:05:55,061 So think about it, you're training on, pre-training on almost millions of tabular data. 65 00:05:55,941 --> 00:06:08,421 Synthetic data, but in a tabular data, it has learned a lot about what are the possibilities of how the numbers are, how the trends are shifting, what kind of trends are even possible in a data and things like that. 66 00:06:08,661 --> 00:06:11,701 It gets some kind of an insight from it in some sense, right? 67 00:06:12,021 --> 00:06:18,661 So, in the same idea, you can think of it that once you have this pre-trained model. 68 00:06:19,381 --> 00:06:33,381 be it for, GPT, for medical application, for finance application, for agriculture application, and whatnot, we can, use the foundation model to, make predictions 0 short or probably few short in some sense. 69 00:06:34,061 --> 00:06:44,501 Okay, so just to, bring the idea of what 0 short or few short means, just like, in ChatGPT, you say, here are some examples of, 70 00:06:45,301 --> 00:06:46,901 how I want the response to be. 71 00:06:47,301 --> 00:06:50,821 That is 0 shot, like few shot where you know you're giving some example. 72 00:06:51,261 --> 00:06:55,461 And 0 shot is essentially where you're just saying that you know this is what it is, give me an output. 73 00:06:55,941 --> 00:06:58,501 So in both the cases you can work with it. 74 00:07:00,181 --> 00:07:11,701 So to summarize all the you know unique you know benefits of using a Tableau foundation model, one foundation model it helps us in you know 75 00:07:12,301 --> 00:07:13,701 reducing the training time. 76 00:07:13,701 --> 00:07:16,101 You don't have to train your own models. 77 00:07:16,101 --> 00:07:18,021 It makes it robust and things like that. 78 00:07:19,141 --> 00:07:24,141 And then it works in especially the regime of low data. 79 00:07:25,141 --> 00:07:30,421 Earlier, if you have to train your own model, then the question comes, you know, how much data do I need? 80 00:07:30,741 --> 00:07:32,861 How many data points should I collect? 81 00:07:32,861 --> 00:07:38,621 And it usually goes into thousands, you know, probably more than thousands as well, millions and all. 82 00:07:39,141 --> 00:07:40,781 But with 83 00:07:41,381 --> 00:07:50,901 these kinds of tabular foundation models, maybe perhaps you just need 20 examples, 30 examples of data, and perhaps you'll get much better results than what you would have imagined. 84 00:07:51,381 --> 00:07:58,581 So that's one particular benefit that I see, particularly in manufacturing, where collecting data is very difficult. 85 00:08:00,661 --> 00:08:07,701 That is 1 particular reason why tabular foundation models are very impactful. 86 00:08:08,261 --> 00:08:11,461 The third is, as I said, data imputation and all, right? 87 00:08:11,461 --> 00:08:14,701 So as I mentioned, you don't have to clean up the data. 88 00:08:14,741 --> 00:08:19,901 You don't have to do any pre-processing per se before you feed the data. 89 00:08:19,941 --> 00:08:24,341 It can understand even if it is noisy, even if the data is missing and all. 90 00:08:25,301 --> 00:08:29,501 And then the 4th is domain agnostic, as I already mentioned. 91 00:08:29,741 --> 00:08:34,261 And then the fifth part is think about it, and that's 92 00:08:35,541 --> 00:08:44,341 Tabular intelligence in itself is something that's being developed in the recent times, and a lot of industries are investing on it. 93 00:08:45,221 --> 00:08:54,581 In fact, Amazon has its own tabular foundation model called Mitra, and Prior Labs is a startup which is working on tabular foundation models and all. 94 00:08:55,061 --> 00:09:05,221 And then there are other startups which are doing a lot of work in building zero-shot and few-shot models, which help us in doing a lot of tabular intelligence. 95 00:09:06,101 --> 00:09:09,781 I do have some demos I can show you to ascend if time permits. 96 00:09:12,821 --> 00:09:36,501 So, to give you an idea of, course there is a whole idea of how the model is trained and other things, but or the architecture of it and all, but the idea is that what you feed the model is your X train, which is your inputs for your training data and or whatever you call as a small set of data as a context that you're providing. 97 00:09:37,061 --> 00:09:40,821 And then the Y train is what is the labels that are supposed to be. 98 00:09:41,301 --> 00:09:46,421 And then X test is the samples, the inputs for which you want to make prediction. 99 00:09:47,141 --> 00:09:49,541 So just like you say here are a few examples. 100 00:09:49,781 --> 00:09:51,061 These are X, these are Y. 101 00:09:51,421 --> 00:09:56,021 And you're asking what is the prediction on a few set of X that you have. 102 00:09:56,421 --> 00:09:58,741 You know the inputs and you want to know what the output is. 103 00:10:00,061 --> 00:10:02,981 Just like you know GPT is a transformer based model. 104 00:10:03,781 --> 00:10:06,341 This is also a transformer-based model. 105 00:10:06,661 --> 00:10:12,021 Tap PFN V2 is one of the Tableau Foundation models that exist right now. 106 00:10:12,581 --> 00:10:16,741 Apart from that, there are many others like Tab ICL, Tab DPT, and all. 107 00:10:17,061 --> 00:10:23,461 These are all Tableau Foundation models, but the most famous one is Tap PFN, Prior Fitted Networks. 108 00:10:23,461 --> 00:10:28,421 It's created by the startup called Prior Labs, based out of Germany. 109 00:10:29,141 --> 00:10:32,261 I think they also have an office in New York now. 110 00:10:32,741 --> 00:10:36,501 So these Tap PFN models are also based on Transformer. 111 00:10:36,501 --> 00:10:40,821 They use the same self-attention across the rows to create a context. 112 00:10:40,821 --> 00:10:55,381 And right now, you know, earlier when we started working on Tap PFN and all, the context window was around 10,000 rows and, you know, 500 features or 500 columns of tabular data. 113 00:10:56,341 --> 00:11:01,621 Now these models can work with almost around 100,000 rows. 114 00:11:02,181 --> 00:11:07,301 and 5000 or 2000 or 5000, one of those number of columns. 115 00:11:08,261 --> 00:11:10,581 Again, that's not a stopping point. 116 00:11:11,061 --> 00:11:16,341 If you are, hitting any of these hurdles, there are ways you can even get around those. 117 00:11:16,901 --> 00:11:23,221 But this is where the current, limits are for these kinds of models. 118 00:11:23,621 --> 00:11:31,861 Just like, when we started out with, in a ChatGPT 3.5 or 4, your context window was, 256,000 tokens or something like that. 119 00:11:32,181 --> 00:11:43,541 But now we are working with 1,000,000 tokens and all, so the same way the context window is something that is right now at this stage, but we are hoping that this expands even further. 120 00:11:45,621 --> 00:11:49,301 Yes, which one? 121 00:11:53,221 --> 00:11:54,901 Yes, so it's a Hallman et al. 122 00:11:54,901 --> 00:11:56,421 is a paper which... 123 00:11:57,461 --> 00:12:04,221 with basically Nature paper which was published on using Tap PFN for scientific applications and all. 124 00:12:04,221 --> 00:12:04,341 Yeah. 125 00:12:07,941 --> 00:12:08,581 Thank you. 126 00:12:08,701 --> 00:12:08,821 Yeah. 127 00:12:10,581 --> 00:12:13,781 I have few examples I wanted to show. 128 00:12:14,501 --> 00:12:17,141 The first example is of a manufacturing machining data. 129 00:12:18,261 --> 00:12:26,341 And you know, there's an extreme example, but you know, I thought, let's start with this extreme example. 130 00:12:28,181 --> 00:12:33,061 This is actually from a paper which came out in around 2010 or something. 131 00:12:34,501 --> 00:12:47,861 This had results of around, the inputs are speed, the cutting velocity, the feed rate for turning operation, and then the depth of cut, and then the nose radius of the tool for inputs. 132 00:12:48,661 --> 00:12:55,141 Output is surface roughness that you get as an output of whatever workpiece you're working on. 133 00:12:55,621 --> 00:13:16,821 And simple manufacturing process, you have 4 input parameters, you have one output with zero shot, just using twenty-two rows, training rows, and using rest of the rest of the samples for testing and all, you get around.938 R-square value. 134 00:13:18,101 --> 00:13:27,221 It's quite good in terms of zero-shot getting this kind of a performance, whereas if you expect, perhaps you train your own model. 135 00:13:28,661 --> 00:13:39,621 When I started working on this kind of a very small data set, the best model I got was around.91,.92, given the same data split and all. 136 00:13:39,621 --> 00:13:42,901 So this was a very good performance when we started out. 137 00:13:46,821 --> 00:13:52,341 This is 1 example, as I say here, most regressors collapse below 50 samples. 138 00:13:53,461 --> 00:13:59,461 You don't really, how do you even work with 50 samples for any of the machine learning models? 139 00:13:59,461 --> 00:14:09,061 And all is a question that quite often people ask, but this is one example where we have been able to do very, with very few shots, being able to do good prediction. 140 00:14:10,181 --> 00:14:10,661 And then 141 00:14:12,421 --> 00:14:14,821 This is another example of a problem. 142 00:14:16,341 --> 00:14:21,541 We are in Iowa, so we should talk about agriculture to some extent or the other. 143 00:14:21,861 --> 00:14:24,781 So this is agriculture yield prediction. 144 00:14:24,781 --> 00:14:29,461 We published this in AAAI workshop earlier this year. 145 00:14:30,501 --> 00:14:31,941 We worked with three data sets. 146 00:14:33,701 --> 00:14:38,341 The 3 data sets are for soybean in US. 147 00:14:39,461 --> 00:14:42,101 It has about 86,000 samples. 148 00:14:43,461 --> 00:14:49,381 And then we have global from multiple regions and all, around 28,000 samples. 149 00:14:49,381 --> 00:14:57,221 And then one specific data set for European Union, around 8,600 samples. 150 00:14:58,181 --> 00:15:05,621 The inputs for this is, you know, as you can imagine, yield prediction, you need to know what is the kind of weather in that area. 151 00:15:05,941 --> 00:15:07,861 It's aggregated features of weather. 152 00:15:08,421 --> 00:15:12,821 and some crop information and things like that. 153 00:15:13,181 --> 00:15:22,021 And then you also have, so this doesn't have any missing values, but as you can see, this one has about 5 to 13% of missing values. 154 00:15:22,581 --> 00:15:26,981 And this is categorical heavy in terms of the samples and everything. 155 00:15:27,221 --> 00:15:31,621 And very heterogeneous in terms of, because you're working with a very diverse and complete sample. 156 00:15:33,141 --> 00:15:39,861 So you can see large, complete, diverse, complete samples, and then the small but missing samples case as well. 157 00:15:40,501 --> 00:15:44,501 So these are, three varieties of, cases that you can see. 158 00:15:45,461 --> 00:16:00,061 As you can see here, tab PFNV2 with almost zero shot performs much better than, you know, all the machine learning models that we have known all along, like, you know, CatBoost, XGBoost, Random Forest. 159 00:16:00,821 --> 00:16:06,181 If you have worked in machine learning in tabular data for a while, you would have heard of any of these terms quite easily. 160 00:16:06,661 --> 00:16:10,741 And you can see that this performs much better than those 0 shot. 161 00:16:11,301 --> 00:16:21,301 And then there's something called as auto gluon, which is essentially, you know, fine tuning whatever you get from tap PFN on top of it to, you know, essentially make it even better. 162 00:16:21,941 --> 00:16:26,821 So you can, that's essentially the whole story that we have here. 163 00:16:29,541 --> 00:16:49,541 Another example is this global case where you can see that, even with zero shot, we are able to get almost close performance to this, but obviously random forest is doing better in this case, but not that different in 9716 to 9794. 164 00:16:49,781 --> 00:16:56,581 It's not like you have a major difference there, but still something to note in that sense. 165 00:16:58,021 --> 00:17:00,501 The key part is the compute part. 166 00:17:00,901 --> 00:17:10,501 You can get this result in less than a second rather than training a model, preparing the model, and doing all the things that you have to do for training and doing any of these things. 167 00:17:11,781 --> 00:17:26,981 Same way you can see this one is when you have missing samples, random forests and all don't do as good, but you know, Tap BFN too, because of all the, you know, auto imputation and things like that, it does much better than, you know, 168 00:17:27,621 --> 00:17:28,421 all the things. 169 00:17:30,821 --> 00:17:46,341 There are a few things I'm probably need to mention is how the data imputation is impacting the entire thread in general, but in terms of performance, it gives you the.91 instead of all the.93s and.97s that you have seen all along. 170 00:17:47,381 --> 00:17:54,581 But this certainly is an example of how it gets impacted in general. 171 00:17:57,381 --> 00:18:08,661 So, this is about, how we can see that, TAP PF and V2 or Tableau foundation models in general can perform much better than, what you have seen so far. 172 00:18:10,261 --> 00:18:24,661 This, if you want to see in terms of, a different when to use what and all, you can clearly see that, when you have a large and complete data set, you can always, if you have a large data set, you know, you can always argue that, you know, I can always 173 00:18:25,261 --> 00:18:44,581 Perhaps, fine-tune my model, in which case you can go with auto glue on or auto ML type architectures, where you create a model using TAPPFN, but you can always fine-tune it with auto glue on type architectures, and you do much better, but if you have diverse and perhaps complete, then you can either go... 174 00:18:44,981 --> 00:18:50,261 with Tableau Foundation Models, or you can even go with AutoML type architectures. 175 00:18:50,581 --> 00:18:57,621 But if you have small and missing data type scenarios and things like that, going with Tableau Foundation Models helps a lot. 176 00:18:58,741 --> 00:19:06,901 So I think one bottom line that you'll see is, especially if you're running into a low data regime, Tableau Foundation Models certainly win. 177 00:19:07,621 --> 00:19:12,661 Second thing that you'll notice is that, you know, it can work with large data as well. 178 00:19:13,141 --> 00:19:17,221 But you can always improve because you have more data, so you can always do better. 179 00:19:18,021 --> 00:19:29,061 And the other thing is, the bigger picture that you need to understand is, Tableau Foundation models are not going to replace traditional machine learning models in any day. 180 00:19:29,381 --> 00:19:33,461 It can, in terms of, you know, it can be fine-tuned. 181 00:19:33,461 --> 00:19:37,301 Tableau Foundation models can be further fine-tuned using AutoML and all. 182 00:19:37,541 --> 00:19:42,261 But in general, the idea that we are trying to say is that, you know, 183 00:19:44,021 --> 00:19:47,381 Use it to get an initial guess, right? 184 00:19:47,621 --> 00:19:59,221 So it's very quick and you can get responses very quick and you can, use it to work on a bigger picture rather than, just the machine learning model that you're trying to work with. 185 00:19:59,781 --> 00:20:12,701 So think about it that, when we talk about physical AI or any of these things, right, simulations and all, we always say, I don't care about the accuracy of the simulation as long as I'm able to quickly iterate over and go. 186 00:20:12,781 --> 00:20:13,461 move forward, right? 187 00:20:14,021 --> 00:20:17,061 That's how the digital twin, the idea of digital twin and all work. 188 00:20:17,381 --> 00:20:31,461 Same way, you know, if your goal is to not just get some kind of a machine learning model, very perfectly accurate machine learning model, but you want to get some kind of, you know, close to accurate model, and then, you know, you want to quickly iterate and see, you know, what else can I do? 189 00:20:31,461 --> 00:20:33,301 Can I, do I need to add more data? 190 00:20:33,301 --> 00:20:36,501 Can I, do I need to bring more other data features and things like that? 191 00:20:36,741 --> 00:20:42,341 I don't want to sit on, you know, keep on training a model when I don't even know whether that model is really what I want to train. 192 00:20:42,741 --> 00:20:45,461 or is the data is the problem or what is the problem, right? 193 00:20:45,821 --> 00:20:56,661 Quite often than not, what I've seen when working with different industries is that, there is data which you need to improve on and you need to also improve on the model. 194 00:20:57,061 --> 00:21:05,861 But this at least helps me in, focusing on the model, on the data, because I know that the model can do as best as what I want in some sense. 195 00:21:07,301 --> 00:21:09,381 There's another example here. 196 00:21:10,981 --> 00:21:29,061 Vehicle sensor data, it's like, what you get from a CAN bus, the sensor data from a CAN bus to essentially, in this case, it's for, large combines to essentially detect some kind of, information of soil moisture or different things that you can get. 197 00:21:31,381 --> 00:21:34,661 Sorry, correct? 198 00:21:35,141 --> 00:21:38,501 Yes, so it is from that. 199 00:21:39,461 --> 00:21:57,941 Again, for the sake of, anonymity, I'm not providing what combined what data and other things, but the idea is we had about, 8 features of, canvas signals aggregated by different unit IDs of experiments that we were doing. 200 00:21:58,341 --> 00:22:07,541 And then it's again a tableau regression model where you're essentially trying to understand in real life, there's a lot of, this is 201 00:22:07,941 --> 00:22:22,661 One of the most noisy data that I've ever seen, it has all kinds of heterogeneity in terms of the inputs going all the way to NANDs, but very high numbers at the same time. 202 00:22:23,221 --> 00:22:27,301 had very low numbers in terms of 10 to the power minus 8 and things like that. 203 00:22:27,621 --> 00:22:35,621 And then it also had a lot of missing values, cases where the sensor was not really robust. 204 00:22:35,661 --> 00:22:42,501 We don't even know whether the sensor can be reliable or not, or should I even use that data or not, and things like that. 205 00:22:42,501 --> 00:22:46,941 And there's no real linear structure that you can work with. 206 00:22:46,941 --> 00:22:50,901 So this is as real as it could get in terms of the data set that you can see. 207 00:22:52,261 --> 00:23:02,341 Here again, you can see all the models kind of give up when this did much better than the rest of them. 208 00:23:02,741 --> 00:23:07,301 Again, of course, you can say it's not that different, 878288. 209 00:23:07,541 --> 00:23:15,461 It's not that different, but the key part is we were able to do this in less than a day. 210 00:23:15,941 --> 00:23:20,661 So we could at least understand what is the data issues, what are the different things. 211 00:23:21,061 --> 00:23:29,701 And we could go ahead and do other things that we wanted to do, because this model is not the only thing that was stopping us. 212 00:23:30,181 --> 00:23:41,701 We wanted to use this model to go build something else for the sensor to improve the sensor, understand what sensors do we need to replace, and things like that. 213 00:23:43,381 --> 00:23:49,621 One thing that you'll see is, you know, especially if you are using the number of samples that you're using, right? 214 00:23:49,941 --> 00:23:50,821 So you can see... 215 00:23:51,461 --> 00:24:00,981 If you're using 10% of the samples, then you get 0.84 type correlation, but if you go all the way to using 50%, you get 0.88. 216 00:24:01,461 --> 00:24:10,581 But you can further keep increasing and see what happens, but in most of the cases, it doesn't do that well after that. 217 00:24:11,061 --> 00:24:15,781 So I think 0.882, and I think it's more or less stuck over there. 218 00:24:15,861 --> 00:24:17,381 It doesn't go further from there. 219 00:24:18,421 --> 00:24:19,221 And 220 00:24:21,061 --> 00:24:31,221 But this, I think one thing I wanted to talk about is, how the industry is moving in terms of different things. 221 00:24:31,221 --> 00:24:39,221 So far, what we have seen is in terms of, giving a data, making the prediction and things like that. 222 00:24:40,981 --> 00:24:46,581 The question that I think mostly all of you may have is, okay, what do I do with it? 223 00:24:47,781 --> 00:24:49,701 How does it matter to me? 224 00:24:50,261 --> 00:24:56,421 And that's where, the idea of other models that I was talking about, like Kumo AI is 1 model. 225 00:24:56,741 --> 00:25:01,141 It's a relational foundation model built on top of a Tableau foundation model. 226 00:25:01,461 --> 00:25:06,101 So think about this as, you know, it works with multiple tables. 227 00:25:06,501 --> 00:25:12,741 It understands the relation between them and tries to use that to essentially have a conversation with you. 228 00:25:12,741 --> 00:25:14,421 can ask questions in terms of, you know, 229 00:25:15,501 --> 00:25:17,221 What are the insights on this? 230 00:25:17,221 --> 00:25:19,941 Then it will essentially identify the relations of all of them. 231 00:25:20,261 --> 00:25:26,661 You don't need to flatten the data of multiple tables together to essentially get one big table and then work with it. 232 00:25:27,141 --> 00:25:40,661 So this kind of a relational foundation model is something that people are using now, especially in DoorDash, Snowflake, and all to understand what are the relations, how do I understand the insights of them, and then go from there. 233 00:25:41,621 --> 00:25:50,261 The other models like AWS has, Mitra on top of that, I think anyone of here, anyone here has heard of Amazon Quick? 234 00:25:52,181 --> 00:26:03,141 So Amazon Quick or AWS Quick is one, another dashboard type platform which has these kinds of features of, you know, having a conversation based on a data set. 235 00:26:04,101 --> 00:26:06,901 You can have conversations based on tabular data set. 236 00:26:06,901 --> 00:26:12,901 You can connect S3 buckets and then directly work with it and have some kind of conversations with there and all. 237 00:26:13,541 --> 00:26:16,381 So that's something that I've seen people do quite a lot. 238 00:26:16,381 --> 00:26:22,901 And again, there are other major players that you can think of in this space which are doing something similar to this. 239 00:26:24,901 --> 00:26:30,261 If you ask me where the future is in some sense, you can think of 240 00:26:31,381 --> 00:26:40,501 Now, obviously, tablet intelligence is something that we have seen quite a bit in terms of how we can use in different spaces. 241 00:26:40,501 --> 00:26:46,581 I've covered agriculture, manufacturing, and autonomous vehicles and things like that. 242 00:26:46,901 --> 00:26:54,741 But you can use it for other applications as well, and medical application, FinOps, and a lot of applications have these things. 243 00:26:56,061 --> 00:27:05,301 Again, the other thing is missing values you usually try to imputate and do something on your own, but here you are using autoimputation and things like that. 244 00:27:05,301 --> 00:27:07,701 And that is something that helps us a lot. 245 00:27:08,181 --> 00:27:16,261 And perhaps that can help us in understanding probably that, maybe missing values are not really, a bug. 246 00:27:16,741 --> 00:27:20,301 It's perhaps something deeper inside that you can get from those things. 247 00:27:20,821 --> 00:27:22,901 And quite often than not, we realize that, you know, 248 00:27:23,701 --> 00:27:28,661 LLMs or AI models can understand data differently from what we do. 249 00:27:29,221 --> 00:27:32,661 So perhaps when we get a different insight than what we have seen so far. 250 00:27:34,661 --> 00:27:42,501 And then I think one thing that is a relief for us is that perhaps you don't need a lot of data. 251 00:27:43,781 --> 00:27:48,501 So far we thought we need to collect a lot of data to train our models and do things. 252 00:27:48,981 --> 00:27:50,901 But perhaps we don't need a lot of data. 253 00:27:51,221 --> 00:27:57,141 We just need few samples, few hundreds or even thousands or even probably a million max. 254 00:27:57,141 --> 00:28:05,861 But you don't need a lot of data to start training your own models or using your own models for doing tabular intelligence particularly. 255 00:28:06,901 --> 00:28:12,021 So with this, since I have some time, I can quickly show you a demo. 256 00:28:13,781 --> 00:28:17,941 But before I go there, are there any questions that I can answer for you? 257 00:28:24,981 --> 00:28:28,461 Edge AI will be a great example where to use this stuff, right? 258 00:28:28,981 --> 00:28:29,221 Yes. 259 00:28:30,741 --> 00:28:40,661 Edge AI is something, it will be useful, but there's one caveat to understand that, you know, these are all foundation models. 260 00:28:41,301 --> 00:28:49,021 Just as much as you can't put a big llama model in an edge device, you'll have such considerations. 261 00:28:49,021 --> 00:28:51,781 But I think this is relatively easy. 262 00:28:51,781 --> 00:28:58,581 You can use it on your own laptop, so it's not that bad in terms of memory and compute and all. 263 00:29:00,461 --> 00:29:01,541 Any other questions? 264 00:29:06,661 --> 00:29:06,981 Thanks. 265 00:29:06,981 --> 00:29:07,781 I appreciate 266 00:29:09,061 --> 00:29:12,821 My background is in metal cutting, so I appreciate that you had the example on turning. 267 00:29:14,021 --> 00:29:20,261 Sometimes tabulated data has a different purpose for why it was constructed. 268 00:29:20,501 --> 00:29:30,421 So like your turning example is essentially a set of experimental test results that map out some of the parameter space, but there's also guidelines 269 00:29:31,141 --> 00:29:46,421 in handbooks that are essentially an encoding of knowledge of look-up tables of ranges of parameters to use or look-up tables for roughness that was achieved under certain conditions. 270 00:29:47,221 --> 00:29:49,141 And sometimes it's a different thing. 271 00:29:49,141 --> 00:29:54,421 It's A look-up table like maybe material properties. 272 00:29:54,501 --> 00:29:59,221 So different metals have different stiffnesses, yield strengths, and ultimate tensile strengths. 273 00:29:59,621 --> 00:30:00,181 And 274 00:30:00,621 --> 00:30:14,661 I'm wondering about the role of the tabular intelligence in the context of a combination of the purpose for which the tabulated data was created and the purpose for which the user is trying to use it. 275 00:30:15,541 --> 00:30:18,821 Yeah, that's an excellent question, right? 276 00:30:18,821 --> 00:30:29,221 So I think that's very close to what I was talking about, the Kumo AI part that, you know, perhaps you may have a lot of large database, right? 277 00:30:30,181 --> 00:30:43,261 You may have, as you mentioned, different material properties, different material manufacturing conditions, and even you may have a database of multiple manufacturing conditions like turning, cutting, milling, and all. 278 00:30:43,541 --> 00:30:48,101 You can have a lot of conditions which can all be part of the same database. 279 00:30:48,901 --> 00:31:00,421 But you can essentially, instead of you writing a particular lookup table or a SQL query, say that this is the data that I want, and then perhaps have some kind of an insight from it. 280 00:31:00,741 --> 00:31:12,701 You could say in a natural language that, hey, I want to find out what are the, just like, you go to perhaps your bank account now and say, I want to know what are the trends of... 281 00:31:12,781 --> 00:31:22,581 my last one year of purchases I've had and things like that, then it will essentially filter out the data that is relevant to it and then provide you some insights from it. 282 00:31:22,901 --> 00:31:25,061 So think about it in that perspective. 283 00:31:25,061 --> 00:31:38,501 So it can essentially do that relational database, understand the relation of multiple data or even filter out the data using a SQL query or something and give you something which is more relevant to what you want. 284 00:31:39,221 --> 00:31:42,021 But again, the key thing is 285 00:31:43,221 --> 00:31:53,701 to know in terms of what data sets exist with you so that you can have that kind of a relational graph built in so that you can actually do something like that. 286 00:31:54,821 --> 00:31:55,541 Does it make sense? 287 00:31:58,381 --> 00:31:59,221 Any other questions? 288 00:31:59,501 --> 00:32:02,741 Thank you, sir. 289 00:32:03,781 --> 00:32:08,181 So I got a question about the low data usage for training the model. 290 00:32:08,501 --> 00:32:08,661 Yeah. 291 00:32:10,261 --> 00:32:28,341 To understand that you don't need as much data to train the model, but if there is inherent bias in the amount of existing data that you're using for training, how does it help with extrapolating it for something that's not there in the data? 292 00:32:29,141 --> 00:32:30,901 So example, right? 293 00:32:30,901 --> 00:32:33,061 So we have temperature data. 294 00:32:34,101 --> 00:32:38,901 All my temperature data is around, say, 100 degrees Fahrenheit. 295 00:32:40,501 --> 00:32:44,901 but there's only a few points that are, I could say, 300 Fahrenheit. 296 00:32:46,661 --> 00:32:51,861 I understand it works on low data, but I don't have enough data for 300 Fahrenheit. 297 00:32:52,741 --> 00:33:04,901 Would it still be able to do predictions correctly with lesser data, or do we have to mash the data in the beginning itself so that there is good spread of it? 298 00:33:05,701 --> 00:33:07,621 Right, so it's a great question, right? 299 00:33:07,621 --> 00:33:09,221 So, yes, there will be some 300 00:33:09,621 --> 00:33:15,701 Bias with the starting data that you start or the data that you're starting with, right? 301 00:33:15,701 --> 00:33:31,061 So, if you're saying that you're only going to start with, say, all 100 degree and then probably one or two samples of 300 Fahrenheit, maybe expecting to get some good results with 300 Fahrenheit may be an over expectation over there. 302 00:33:31,421 --> 00:33:34,501 Obviously, the bias is built in the model per se, because... 303 00:33:35,181 --> 00:33:44,341 What it is doing is it's seeing some kind of a, relation or a trend within the data and, saying it doesn't understand that it's a temperature. 304 00:33:44,341 --> 00:33:48,821 It doesn't even understand that, you know, from one temperature to another temperature regime, something is changing. 305 00:33:49,141 --> 00:34:01,861 Just as much as, you know, ChatGPT, if you give a bunch of things and ask something as an output, it may not even do because it doesn't understand the connection between, you know, multiple files that you have provided and, you know, what is it that you're asking as an output. 306 00:34:02,741 --> 00:34:13,141 So same way, that extrapolation capability is certainly going to be dependent on the bias on the data that you're providing in some sense. 307 00:34:13,461 --> 00:34:18,581 If you provide a very clean data of, fully balanced data, then it may do much better. 308 00:34:19,701 --> 00:34:30,981 I think the question that we should probably look for is, you know, the way to rephrase it is, given the data, 309 00:34:31,941 --> 00:34:40,901 The best model performance that you could get in very, less time is going to be what you get from Tableau Foundation models. 310 00:34:41,701 --> 00:34:49,461 You could probably perhaps invest more energy to slightly move it by a little bit, but data is the king ultimately. 311 00:34:49,461 --> 00:34:51,421 You know, you need to probably work on the data. 312 00:34:51,421 --> 00:35:00,461 And that's where you'll probably, you know, if you are starting out, you see that, you know, there are a bunch of 300s and, you know, the rest of the data is in hundreds. 313 00:35:00,461 --> 00:35:01,781 You see, this is the best performance. 314 00:35:01,901 --> 00:35:04,021 performance you can get, is it sufficient? 315 00:35:04,741 --> 00:35:08,421 Maybe it's sufficient because you're probably doing anomaly detection. 316 00:35:08,661 --> 00:35:13,101 It doesn't matter whether it's predicting, thinking it is 300 or thinking it's 150. 317 00:35:13,461 --> 00:35:15,141 It's just for predicting anomaly. 318 00:35:15,141 --> 00:35:16,501 It's anomalous, so it's good. 319 00:35:16,741 --> 00:35:18,661 So you don't need perhaps more data. 320 00:35:18,981 --> 00:35:24,741 But if you are making more clear prediction of, you know, some specific trend of, you know, how... 321 00:35:25,621 --> 00:35:29,621 metal manufacturing processes from 100 degrees to 300 degrees. 322 00:35:29,861 --> 00:35:35,621 There's a complete difference on the formability, the material properties and everything change quite drastically between these two regimes. 323 00:35:35,621 --> 00:35:41,301 Then in that case, perhaps you may need some more data to collect in the rest of the regime. 324 00:35:41,301 --> 00:35:47,621 So depending on the data, but at the same time, it gives you a good quick start for you to go from there, basically. 325 00:35:48,981 --> 00:35:51,661 Thank you. 326 00:35:51,661 --> 00:35:52,581 Any other questions? 327 00:35:55,701 --> 00:35:56,101 All right. 328 00:35:56,821 --> 00:36:03,421 Then if there are no further questions, I mean, you can, if you have questions, I can answer them later on as well. 329 00:36:03,421 --> 00:36:07,301 But I still want to see if I can show you a quick demo. 330 00:36:11,861 --> 00:36:17,461 So there are, you know, I am using two examples for a demo. 331 00:36:17,541 --> 00:36:19,861 One is from prior labs dot AI. 332 00:36:20,581 --> 00:36:29,541 That's the startup which essentially runs or built the model, the TAP PFN V2 model that I was talking about. 333 00:36:30,181 --> 00:36:41,941 And you can see that you can actually upload your own data set and play around with it and do things, especially, you know, on any, like either this model or previous models and things like that. 334 00:36:42,661 --> 00:36:47,621 And in this case, I just chose one of the samples data. 335 00:36:47,941 --> 00:36:51,141 As you can see, there are a lot of samples that are there already. 336 00:36:52,021 --> 00:36:54,661 It can be either sales or it can be industrial. 337 00:36:54,661 --> 00:36:58,181 And you can see that there's a concrete compressive strength data set. 338 00:36:58,261 --> 00:37:00,581 That's what I had loaded earlier. 339 00:37:06,981 --> 00:37:09,781 As you can see, it has a bunch of features. 340 00:37:10,341 --> 00:37:14,581 And then finally, the target and what the prediction is in some sense. 341 00:37:15,701 --> 00:37:29,701 As you can see, TAP PFN in this case for this particular data set, which I can provide the exact metrics, but it gets better performance than even random forest, XGBoost and all. 342 00:37:31,141 --> 00:37:34,261 And linear regression has the maximum error. 343 00:37:34,261 --> 00:37:39,781 Tap PFNV2, 2.5 plus has the minimum error in MSE. 344 00:37:40,661 --> 00:37:43,061 And it provides a simple output. 345 00:37:43,381 --> 00:37:46,821 So this is, and you can easily upload your own data set. 346 00:37:47,301 --> 00:37:57,141 It allows you to directly upload either a CSV file or an Excel file with header rows of around 20 to 40,000 rows. 347 00:37:57,621 --> 00:38:01,061 and including a column on what to predict and all. 348 00:38:02,181 --> 00:38:08,901 So this is a simple, you set interface-based way of how you can do it. 349 00:38:09,301 --> 00:38:16,101 Or if you are more like me who likes to code, then you can always get the code, run it on your own local machine. 350 00:38:16,341 --> 00:38:17,901 You don't want to use that server. 351 00:38:17,901 --> 00:38:19,781 You want to run it on your own local machine. 352 00:38:20,341 --> 00:38:23,781 You can just download the model and then run it in your own machine, local machine. 353 00:38:24,181 --> 00:38:26,261 And that is also equally easy. 354 00:38:26,501 --> 00:38:28,981 You can just access it from here and then run it. 355 00:38:30,741 --> 00:38:34,661 This is just how to run just the Tabular Foundation model alone. 356 00:38:35,061 --> 00:38:41,381 And there is another example, which is the Kumo RFM that I was talking about. 357 00:38:41,381 --> 00:38:49,221 So this is an example of, you can see in the Kumo RFM, they already have few data sets here. 358 00:38:50,701 --> 00:38:52,501 One of them is e-commerce. 359 00:38:53,741 --> 00:39:12,101 Where they have data on returns, views, items, orders, and users, and all, or the other data sets like insurance, F1 racing, and all, or you can even upload your own data set or link it to your Amazon S3 buckets or Snowflake for that matter as well. 360 00:39:12,661 --> 00:39:19,541 And once you have, either you can infer the schema or you can actually write down the schema as well. 361 00:39:19,861 --> 00:39:21,701 That's an option that you can do. 362 00:39:21,941 --> 00:39:35,381 So once you provide the schema and all that information, it will create a graph, something like this to, you know, come up with the entire, you know, the whole idea of how each data is related to other and table is related to other. 363 00:39:35,941 --> 00:39:41,301 And once you have all of this ready, you know, you can, once you have the data, you can always go here. 364 00:39:42,901 --> 00:39:47,861 And all you have to do is, you have to select which table are you working with. 365 00:39:48,421 --> 00:39:56,981 So you say you're saying e-commerce, then it will say how many orders will each user have in next 30 days. 366 00:39:57,301 --> 00:39:59,461 And that's a question that you're asking. 367 00:39:59,461 --> 00:40:01,061 And then it will analyze your question. 368 00:40:02,101 --> 00:40:11,741 It will essentially, if you can see here, it is making a query of what in a SQL query, the product, in terms of a predict query language, they say. 369 00:40:12,981 --> 00:40:23,701 where you're saying, we are predicting the orders and for each user, and then it essentially finds out what is a SQL query that it needs to run and create a table. 370 00:40:23,781 --> 00:40:28,101 And then once it has the table, it will make a prediction based on that. 371 00:40:29,541 --> 00:40:33,221 And you can further ask more questions and, you know, have a conversation in some sense. 372 00:40:34,821 --> 00:40:45,301 Just wanted to show you these two examples of, how you can use this to create tabular data and in a tabular foundation, use this for inference. 373 00:40:45,461 --> 00:40:45,861 Yes, Vijay. 374 00:40:47,221 --> 00:40:49,301 What was the first one? 375 00:40:49,941 --> 00:40:52,181 What was the first tool that you showed? 376 00:40:53,581 --> 00:40:54,981 It's called Prior Labs. 377 00:40:54,981 --> 00:40:57,941 Prior, like, Tap PFN is the model. 378 00:40:58,261 --> 00:41:01,461 Prior Labs is the startup which actually trained that model. 379 00:41:02,181 --> 00:41:02,501 Thank you. 380 00:41:06,741 --> 00:41:08,741 Yes, it's free, of course. 381 00:41:08,741 --> 00:41:10,621 I have not paid a single cent so far to them. 382 00:41:10,621 --> 00:41:21,781 Do you have a link to the Yes, and my slides will be there, and they have that will have the link, so you should be able to access it from there as well. 383 00:41:21,781 --> 00:41:22,021 So, yeah. 384 00:41:23,221 --> 00:41:25,941 I was going to ask a question in those regards, too, about favorite tools. 385 00:41:25,941 --> 00:41:27,141 Obviously, this is one of them. 386 00:41:27,141 --> 00:41:30,501 Any other favorite tools based on benefits that they might have over this? 387 00:41:32,181 --> 00:41:46,181 So, the thing is, this area is quite in as any AI models in space, ChatGPT, like 5.5, and then Opus, Cloud Opus, they're all fighting with each other, same way there are. 388 00:41:46,661 --> 00:41:55,061 So, this is when we started working on it, Tap PFN was the V1, and then we had Tap by CL, and then Tap DPT. 389 00:41:55,061 --> 00:41:57,141 So, these are the three major ones. 390 00:41:57,781 --> 00:42:06,421 Tap PFN, ICL, ICL is in context learning, and then DPT is, I think, some transformer predictive transformer. 391 00:42:06,821 --> 00:42:12,501 I don't know what the D is on top of my head, so these three models have been fairly... 392 00:42:12,781 --> 00:42:13,381 really good. 393 00:42:13,901 --> 00:42:17,301 And then Amazon Smithra is the other one which came out very recently. 394 00:42:18,421 --> 00:42:23,061 Some of these four, three or four are the ones which are right now are doing really good. 395 00:42:24,181 --> 00:42:32,501 If you ask me which one is best so far, I think prior labs, the PFN is well tested in so many broad areas. 396 00:42:33,141 --> 00:42:36,341 You know, Hitachi and so many companies have already used it. 397 00:42:36,581 --> 00:42:38,981 I've myself worked with several industries to, you know, 398 00:42:39,621 --> 00:42:42,741 help them use tabular models for their problems and all. 399 00:42:42,981 --> 00:42:45,061 So Tab BFN is the first go-to. 400 00:42:45,301 --> 00:42:49,141 If not, you can go to Tab ICL and Mitra is the third one. 401 00:42:49,421 --> 00:42:51,861 And that's the rating if you want, if you ask me today. 402 00:42:52,421 --> 00:42:53,501 Tomorrow, I don't know. 403 00:42:55,621 --> 00:42:55,941 Yes. 404 00:42:55,941 --> 00:43:01,421 Is it safe to use like company code right now or are they training models off the data you give them? 405 00:43:01,421 --> 00:43:05,061 So your question is, it safe to use it on company code and all? 406 00:43:05,061 --> 00:43:07,181 Is that company data? 407 00:43:07,181 --> 00:43:08,221 Yes, absolutely. 408 00:43:08,221 --> 00:43:08,901 Because 409 00:43:10,181 --> 00:43:14,341 Especially when I work with industries, I do not use the user interface of this. 410 00:43:14,741 --> 00:43:18,101 I literally have the trained, like download the trained model. 411 00:43:18,581 --> 00:43:24,901 It is so small that it even runs, I can do inference on my Mac, that one, so and it runs on that. 412 00:43:25,221 --> 00:43:25,381 Yeah. 413 00:43:26,501 --> 00:43:28,101 I think we have time for one more question. 414 00:43:28,181 --> 00:43:28,581 Yes. 415 00:43:30,501 --> 00:43:32,581 You say you download the code, you'd use it. 416 00:43:33,261 --> 00:43:37,461 And the thing is, I've been asking for Macs because I think 417 00:43:38,181 --> 00:43:47,621 I mean, right now, because the Nvidia, you're fighting gamers for it, too, and the price is crazy, and so it's like we say it's like get the Macs, right? 418 00:43:48,221 --> 00:43:48,341 Yeah. 419 00:43:48,341 --> 00:43:51,461 Do you feel that that's like a policy to go? 420 00:43:52,061 --> 00:43:53,301 Well, not really. 421 00:43:54,221 --> 00:43:59,061 I mean, it so happened that I'm using a Mac and it's working very good. 422 00:44:00,101 --> 00:44:05,941 I did my PhD using GP computing and all, so yes, I understand that you know you would go with that. 423 00:44:06,581 --> 00:44:14,821 Perhaps the other alternative is you can train these kinds of models are easily accessible from even Google Colab or things like that. 424 00:44:14,901 --> 00:44:17,181 So that's another way you can quickly train the model. 425 00:44:17,181 --> 00:44:20,781 It's going to be very easy to do it in Google Colab as well. 426 00:44:23,461 --> 00:44:27,781 I agree, but Google Colab is a temporary instance. 427 00:44:28,021 --> 00:44:31,301 You will probably train the model then download it to your local. 428 00:44:32,981 --> 00:44:33,301 Yes. 429 00:44:34,421 --> 00:44:35,221 Any other questions? 430 00:44:36,821 --> 00:44:38,501 I think we're right at our time there. 431 00:44:38,501 --> 00:44:41,061 So everybody, please join me in thanking Dr. 432 00:44:41,061 --> 00:44:41,981 Blue for his presentation.