028 - Easy entry into the world of AI in fire with MZ Naser
Have you ever been fascinated by the capabilities of AI? Did you wonder how the heck can an algorithm beat humans in repetitive tasks? Or make multi-level correlations that we would never be able to figure out? I was as well. And I felt the urge to learn more about this technology, in a way to not be left out when everyone plays with their new toys... But at the same time, I felt this feeling of overwhelm and confusion about this technology. What exactly is it, where to start... Then the wall of multiple choices to take - am I even trying supervised or unsupervised learning? Is my problem a regression or classification? I won't lie, it's hard already, and I have not even really started yet.
And then comes him. Dressed in white (just kidding). MZ Naser.
MZ is not only a genius who seems to have figured it out in the world of fire, but he is also documenting every step of his path in research papers. More to that, he also wrote a bunch of entry-level papers, and a review paper summarizing the basics and explaining the core concepts. Wow, what a service to the community! Please join me in this discussion in MZ, where he literally walks me through the fascinating world of AI in fire and explains where to start.
At this point of the show notes, I would list you a bunch of papers and relevant resources.
(Update: originally there was just link to MZ’s site, but as this got published you totally need to start with this paper: https://www.readcube.com/articles/10.1007/s10694-021-01210-1 )
MZ Naser is so nice, he has a website where all of this is summarized and kept updated! If I started listing the resources, I would do you a disservice... You have to test what he has out there.
https://www.mznaser.com/
And also, please connect with MZ at his Twitter and LinkedIn
----
The Fire Science Show is produced by the Fire Science Media in collaboration with OFR Consultants. Thank you to the podcast sponsor for their continuous support towards our mission.
Available Results
Generated results are saved to the knowledge database for reuse and search.
Extract Knowledge
Pick what you want extracted first. Model, scope, and chapter options appear after a template is selected.
Transcript
WEBVTT 00:00:00.540 --> 00:00:01.169 Hello, everybody. 00:00:01.169 --> 00:00:03.359 Welcome to Fire Science Show session 28. 00:00:03.839 --> 00:00:07.769 Today we'll be discussing one of my favorite topics in all of the fire science. 00:00:07.769 --> 00:00:09.929 And that is the use of artificial intelligence. 00:00:10.044 --> 00:00:15.833 I consider it one of my favorites because it's something that I would really, really love to learn myself. 00:00:16.089 --> 00:00:21.969 And I'm exploiting this podcast to bring me the best guests that can explain it to me a bit more. 00:00:22.449 --> 00:00:26.640 Honestly, I was quite confused where to start with, , with all of this. 00:00:26.690 --> 00:00:33.320 And after the discussion, as you will hear in this episode, I have a little bit better idea of where to start and how to start. 00:00:33.890 --> 00:00:38.810 And actually I should go on this journey because I'm absolutely convinced it's worth it. 00:00:39.530 --> 00:00:42.710 Today with me I have one of the young leaders of fire safety. 00:00:43.299 --> 00:00:54.820 He's an author of literally countless papers, on AI, in the fire science, I'm actually astonished by the amount and quality of work he's putting through and, publishing. 00:00:55.130 --> 00:00:57.399 I'm really, really admiring him for this. 00:00:57.939 --> 00:01:00.939 Um, he's a professor at Clemson University. 00:01:01.280 --> 00:01:06.003 Yeah, let's just jump into not prolong this because, you want to hear what's after the intro. 00:01:06.003 --> 00:01:34.436 Please welcome MZ Nasser and let's jump into the world of AI and fire science! Hello everybody. 00:01:34.436 --> 00:01:35.846 Welcome to Fire Science Show. 00:01:35.965 --> 00:01:39.956 Today, I'm here with professor MZ Naser from Clemson university. 00:01:40.075 --> 00:01:41.516 Hey Naser great to have you here. 00:01:41.695 --> 00:01:42.266 How about you? 00:01:42.475 --> 00:01:43.016 How's it going? 00:01:43.019 --> 00:01:44.099 I'm fantastic. 00:01:44.099 --> 00:01:45.929 I'm about to learn so much about AI. 00:01:45.929 --> 00:01:49.025 I'm really happy to, or that, Hope so the house. 00:01:49.210 --> 00:01:54.834 I've invited you here because you are a rising star in our industry. 00:01:54.834 --> 00:02:12.281 And you're probably one of very few people who has a good clue about how AI is working and how can we use it, in engineering, actually, there's not so many of them, AI luminaries in our community and, through your papers, which are actually very educative. 00:02:12.311 --> 00:02:17.260 They're not like flashing out you see how advanced algorithms I can use? 00:02:17.260 --> 00:02:25.330 You, you publish a lot of like introductory level AI papers, review papers, and they appreciate that so much. 00:02:25.704 --> 00:02:30.300 What puts you on this pathway to, use computers to enhance your learning? 00:02:30.781 --> 00:02:32.550 This is a very good question, actually. 00:02:32.911 --> 00:02:39.591 So the first time I learned about AI at was, sometime in 2012 or 2013, I was taking a transportation course. 00:02:40.121 --> 00:02:48.021 And in that course, the professor was discussing how we can use AI to organize traffic, traffic, lights, synchronize, all different types of Mm. 00:02:48.111 --> 00:02:48.771 infrastructure. 00:02:49.342 --> 00:03:06.316 And then I kind of did like a very short paper at the time on, fire, on, I get lost because, you know, once you go to your PhD, you focus on experimentations, every simulation of the, uh, you know, those things, then you kind of forget the AI and once I was done, was trying to find the faculty job. 00:03:06.316 --> 00:03:11.445 And as you know, fire tech experiments are very, very expensive and you have to have lab and equipment. 00:03:11.445 --> 00:03:16.346 As you know, you have a massive lab here, in my case, it was very hard to develop this left. 00:03:16.346 --> 00:03:21.536 So I was, I had to do something and I wanted to do something a little bit different than simulation. 00:03:21.536 --> 00:03:33.526 And so I went back to my road to AI and that's when things clicked back again, From 2012 2013, maybe 2019 things has rapidly changed on the AI front. 00:03:33.705 --> 00:03:34.966 Many, many things has changed. 00:03:34.966 --> 00:03:38.145 We have different algorithms, different training systems, learning systems. 00:03:38.716 --> 00:03:41.489 So I had to everything else at that time. 00:03:41.908 --> 00:03:44.729 And hence, some of my papers are just like what you mentioned. 00:03:44.729 --> 00:03:49.929 They are very on the same level because as I was writing them, I was also learning. 00:03:50.468 --> 00:03:51.674 I decided Okay. 00:03:51.758 --> 00:03:53.968 a very smooth way because this is how I learn. 00:03:53.979 --> 00:04:01.598 So perhaps it will also be easier to somebody who was as familiar with AI as me at that time to go about with those papers. 00:04:01.854 --> 00:04:03.504 So, so that's a, that's such a cool path. 00:04:03.504 --> 00:04:08.305 So we basically we're documenting your own ways through the world of AI. 00:04:08.694 --> 00:04:09.985 That's uh, that's so cool, man. 00:04:10.495 --> 00:04:24.081 And that also confirms the theory that you just have to be one step ahead from others, uh, to be an expert and, you know, to, to teach you don't have to know everything to, to provide useful guidance. 00:04:24.081 --> 00:04:26.781 And I really appreciate that you are doing that. 00:04:27.411 --> 00:04:29.545 And I assume on your path to the. 00:04:29.971 --> 00:04:38.742 AI, uh, you've stumbled upon the same confusion everyone is stumbling against for me. 00:04:38.742 --> 00:04:42.045 It's like, What the hell is AI after all? 00:04:42.095 --> 00:04:43.475 Can you even define it? 00:04:43.985 --> 00:04:47.045 And then w what kind of AI should I go? 00:04:47.045 --> 00:04:57.024 Because once you start, like digging, you enter this, uh, loophole with hundreds of algorithms, models approaches, and it's really confusing. 00:04:57.295 --> 00:04:58.314 So, yeah. 00:04:58.504 --> 00:04:59.605 How was it for you? 00:04:59.709 --> 00:05:00.666 Yeah, a hundred percent. 00:05:00.666 --> 00:05:01.505 I didn't know. 00:05:01.956 --> 00:05:03.725 four years ago, I didn't really know. 00:05:03.843 --> 00:05:05.403 or I only knew neuron networks. 00:05:05.583 --> 00:05:07.322 This is what I was Okay. 00:05:07.892 --> 00:05:09.062 it prepped on earlier. 00:05:09.375 --> 00:05:14.016 To me, if this was my math, I could solve anything with neural networks because it was in know Hm. 00:05:14.596 --> 00:05:18.495 one tool that you can button the data points, it would run. 00:05:18.495 --> 00:05:24.136 It should give you some kind of a good performance, if not performance on different problems. 00:05:24.675 --> 00:05:29.115 as you mentioned, nowadays, we have all these different types of learning or these of algorithms. 00:05:29.115 --> 00:05:32.326 And the easiest thing in my case was I need to learn. 00:05:32.326 --> 00:05:32.716 I need to learn. 00:05:33.451 --> 00:05:37.560 So I had to go back and see, okay, what are the basics for supervised learning? 00:05:37.560 --> 00:05:39.060 What, what is classification? 00:05:39.060 --> 00:05:45.911 What is regression and , so once you go back to computer science and see those definitions, then you would see, okay, know what? 00:05:45.911 --> 00:05:49.451 Most of the problems we do in fire engineering are really regression. 00:05:49.661 --> 00:05:50.990 You know, we have a phenomena. 00:05:51.446 --> 00:05:54.985 And the outcome of this phenomenon is a number in a fire resistance. 00:05:55.196 --> 00:05:58.586 It could be like, you know, heating, great burning, great, some kind of a number. 00:05:59.185 --> 00:06:14.586 you're out to some kind of a number, then this is a very good chance that you are going to be dealing with a supervised learning problem with a sub a component that's going to be a regression if your output is going to be something like a category, for instance, this column fails or doesn't fail, slap collapse, it doesn't collapse. 00:06:14.946 --> 00:06:16.935 You have charring and you don't have charring. 00:06:17.646 --> 00:06:25.103 instance, this fire is heat, uh, in a ventilated control, you are trying to put the phenomenon into one group, this is classification. 00:06:25.523 --> 00:06:36.862 So once you know the problem, you have to define the algorithm, then you'd say, okay, well now my problem is for instance, regression, what kind of algorithms are there out there now that can solve a regression problem? 00:06:37.360 --> 00:06:42.029 So from there you will go, you'll find hundreds of algorithm that can do the same. 00:06:42.819 --> 00:06:46.509 So the question becomes what, which one of these algorithms I'm going to, I'm going to use. 00:06:46.540 --> 00:06:54.860 And the answer is to be honest, you could potentially use any single algorithm of these and if you have a good database, you will come up with a good answer. 00:06:55.040 --> 00:07:06.870 You would come up with a good prediction to the problem becomes, as you might have guessed is why would I go with algorithm A, instead of algorithm, B or C or D what are the motivation behind these algorithms? 00:07:07.505 --> 00:07:13.386 the answer to this is interesting because this is exactly like saying, shall I use ANSYS or Abacus to solve a problem? 00:07:14.055 --> 00:07:16.565 It's basically which algorithm you're familiar with. 00:07:16.625 --> 00:07:17.406 It's the company. 00:07:17.507 --> 00:07:19.017 To your own experiences. 00:07:19.017 --> 00:07:20.637 In my case, I've always used ANSYS. 00:07:20.668 --> 00:07:23.067 I use Abacus very, very slightly. 00:07:23.067 --> 00:07:26.697 So if you go back to my papers, they're all ANSYS, thing with my algorithms. 00:07:26.697 --> 00:07:29.997 You'll see that the earlier work was heavily towards neural network. 00:07:30.867 --> 00:07:40.257 recently, I've learned more about different algorithms, the modern ones, because now as you know, more than algorithms are almost superior when it comes to trying to for prediction power. 00:07:41.098 --> 00:07:54.160 to be honest, if you have a nice database, if you run, let's say 10 algorithms out of the 10 algorithms, most likely nine of them will give you R square or of 95%, 90%, 85%. 00:07:54.819 --> 00:08:01.343 it's the science is really not in running the algorithm, the sciences what did you learn from this algorithm? 00:08:01.822 --> 00:08:03.382 let's say that you use this algorithm. 00:08:03.382 --> 00:08:05.589 You have a good performance, but how does this. 00:08:05.697 --> 00:08:18.446 Actually advances our science or our knowledge, So to break the first wall for anyone jumping into, , I, from your papers I've learned or supervised, unsupervised and semi-supervised methods. 00:08:18.466 --> 00:08:24.543 And, this seemed like the very first critical choice, uh, one would make, uh, when they enter. 00:08:24.543 --> 00:08:37.009 So could you like try and briefly showcase the differences and, and give these examples, like So supervised learning term supervise means, you know, the inputs you know, the output. 00:08:37.129 --> 00:08:38.159 So everything is being. 00:08:38.784 --> 00:08:43.559 Um, So for instance, let's say that we are trying to figure out if this column is going to fail under fire. 00:08:43.919 --> 00:08:47.009 We know the column, geometry, we know its material properties. 00:08:47.009 --> 00:08:51.600 We know if it's going to be boundary conditions fixed, wing, all of these things, we've done a test. 00:08:51.690 --> 00:08:57.720 So we also know it's fire resistance, or we know when it's going to fail, you know, everything, you know, the inputs, you know the output. 00:08:57.730 --> 00:08:58.620 So this is supervised. 00:08:59.190 --> 00:09:07.472 Let's say now, know, all the inputs, let's say you have a group of, columns, you know, all their inputs, but you don't know when they fail. 00:09:08.269 --> 00:09:15.645 So you would use unsupervised learning and this way the algorithm should cluster or combine the columns that are similar to each other into groups. 00:09:15.645 --> 00:09:22.696 And then the algorithm would say, well, these five columns are group one, these four columns or group two, these four columns are group three. 00:09:22.905 --> 00:09:24.405 You don't know the output. 00:09:24.870 --> 00:09:34.490 You don't know why these are in groups, but if you go back and study the fire test results, you're likely to see that the columns and the group one, maybe they failed within an hour. 00:09:34.809 --> 00:09:35.868 Group two maybe they Okay. 00:09:35.942 --> 00:09:36.753 within two hours. 00:09:37.352 --> 00:09:41.072 unsupervised learning is when you know the inputs, but you don't know the output. 00:09:41.082 --> 00:09:42.543 You don't know what the phenomena is. 00:09:42.602 --> 00:09:43.982 You're just trying to group things together. 00:09:45.062 --> 00:09:47.462 Semi-supervised learning is going to be somewhere in between. 00:09:47.462 --> 00:09:53.822 Sometimes semi-supervised would be something that say that we have, images of columns failing. 00:09:54.452 --> 00:10:00.243 instead of us going by image and saying this column, fails this column doesn't fail. 00:10:00.572 --> 00:10:10.452 We could potentially only label 50 images and the algorithm should be able to label the additional 50 images that we did in labor. 00:10:10.452 --> 00:10:13.302 So this way it has a little bit of knowledge on the inputs outputs. 00:10:13.753 --> 00:10:15.373 doesn't have it for all the database. 00:10:15.972 --> 00:10:19.212 So it's somewhere in between supervised and unsupervised learning. 00:10:19.772 --> 00:10:37.635 So, if you had a, let's say a supervised algorithm with the database on existing columns, and then you come up with a completely new column, the supervise would tell you when it will fail, based on its knowledge, the unsupervised would, it would tell you to which group of columns this one looks more familiar to. 00:10:37.635 --> 00:10:42.706 And the semi-supervised could just continue the task you were doing with the previous columns. 00:10:43.035 --> 00:10:47.655 Was it painting it pink or measuring their moment of inertia or something? 00:10:47.655 --> 00:10:48.105 Okay. 00:10:48.775 --> 00:10:49.676 This seems useful. 00:10:50.123 --> 00:10:50.633 Nice. 00:10:50.663 --> 00:10:53.349 Uh, reminded me when you said, your technician has an AI. 00:10:53.589 --> 00:10:58.678 This is what his algorithm would do, algorithm would recognize noise, or maybe like my mic the hoodie. 00:10:59.158 --> 00:11:12.089 it would label that as this is like noise or this is like not voice because it has seen before through training that this sign of scratching is not really a voice, so you have to take it away. 00:11:12.535 --> 00:11:22.885 let's just jump quickly from enthusiasm to the dangerous region, because you've mentioned it seen, but if has not seen something, it's very unlikely. 00:11:22.885 --> 00:11:24.926 It's going to predict the behavior, right? 00:11:25.166 --> 00:11:31.916 Like if you, if you show it a thousand fires with flashover, it will not know that backdraft my may happen. 00:11:31.916 --> 00:11:32.155 Right. 00:11:32.765 --> 00:11:33.696 And this is the problem. 00:11:33.905 --> 00:11:35.135 This is exactly the problem. 00:11:35.166 --> 00:11:53.105 The problem is when you develop an algorithm and have a good database and you have good performance, the researcher or need to remember that this performance is only valid for your database to go beyond the database is going to be very, very tricky because when you have a database, you're immediately constraint in your algorithm. 00:11:53.115 --> 00:11:55.841 So you have a space of, oh, you have a okay. 00:11:56.015 --> 00:11:56.975 a space of inputs. 00:11:57.446 --> 00:12:01.765 You can possibly collect everything and you can collect some features of the space. 00:12:02.186 --> 00:12:07.166 And then for the algorithm on what it sees is features as the whole space. 00:12:07.346 --> 00:12:10.602 So if you have an additional feature outside of this space. 00:12:11.023 --> 00:12:13.778 It's going to be very hard to give you a correct production. 00:12:13.778 --> 00:12:18.399 Maybe it could sometimes if, if the algorithm or maybe if the problem is simple enough, it could. 00:12:18.879 --> 00:12:23.739 But other than that is going to be very tricky that's very similar to experience, actually. 00:12:24.009 --> 00:12:24.249 Yeah. 00:12:24.349 --> 00:12:29.139 If you experienced a lot of things, you're more likely to predict things. 00:12:29.239 --> 00:12:33.105 That's uh, that was something we share with the machines, I guess. 00:12:33.138 --> 00:12:33.317 Yeah. 00:12:33.317 --> 00:12:43.780 W one thing, yeah, experience, value this a lot when we have experienced, usually like at least us, we have a knowledge of what could happen. 00:12:43.811 --> 00:12:51.051 Like we could see beyond that we have algorithms can't and that's, that's going to be the problem that we're going to be dealing with. 00:12:51.270 --> 00:12:54.380 We can go beyond what we can see beyond the data. 00:12:54.620 --> 00:12:58.490 However, very, very good to see between the lines and we're not. 00:12:59.030 --> 00:13:06.380 So this is how we can compliment both of us, because if you have a complex database for us, it's going to be very hard to visualize for them. 00:13:06.380 --> 00:13:06.890 It's easy. 00:13:06.890 --> 00:13:15.073 If it can see things, and this is why they predict things with high accuracy, but that doesn't mean that this prediction is actually something that's physically correct. 00:13:15.645 --> 00:13:44.566 I had this episode on AI and fire already with Xinyan Huang from Hong Kong Polytechnic University, and Xinyan is doing a lot of crazy things with, , with smoke control fire detection in tunnels and you also mentioned that this, human machine, uh, combination is the most powerful and in a way I had a feeling he would like this AI be a way you could transfer the collective experience of whole industry. 00:13:44.936 --> 00:13:51.316 And that for me, that was such a powerful and beautiful idea that, , so much knowledge is lost between us. 00:13:51.316 --> 00:14:01.546 And if we could have this collective mind helping each other, it would be fun, but it's also seems very difficult from the technical point of view to achieve that. 00:14:01.546 --> 00:14:01.875 Right. 00:14:02.296 --> 00:14:14.298 Because, like, to what extent the, um, structure of the database is also important, like to what extent you can drop scattered data into an algorithm and expect correct results. 00:14:14.817 --> 00:14:14.899 Okay. 00:14:15.533 --> 00:14:17.634 first of all, we're not computer scientists. 00:14:17.634 --> 00:14:20.139 were appliers.Computer Yeah. 00:14:20.224 --> 00:14:24.173 the algorithms, they validate them over multiple databases. 00:14:24.173 --> 00:14:24.953 We just take them. 00:14:24.953 --> 00:14:28.553 And then we do our own little experiment and we have good performance and we think it works. 00:14:29.283 --> 00:14:46.293 second part of the issue as the machine learning we're using now, or the algorithms that we're using now, they're highly data driven or correlations driven which negates the purpose of science and not everything correlates that there is a cause and effect. 00:14:46.803 --> 00:14:48.063 This is why, Okay. 00:14:48.244 --> 00:15:02.494 myself, I'm trying to move away from all this data driven nonsense and go towards like modern algorithms that at least can give you a cause and effect because if you know the cause and effect, but regardless of how much data you have, always get the right answer. 00:15:02.644 --> 00:15:04.384 The goal is to know why this. 00:15:05.283 --> 00:15:12.874 That goes, the issue is not to know seeing this, or I have seen this in 10 experiments and this would happen in the 11th extrovert experiment. 00:15:12.884 --> 00:15:13.724 There is no guaranteed. 00:15:14.203 --> 00:15:15.403 Observations help. 00:15:16.004 --> 00:15:19.844 However, to come up with knowledge, you need to know why, and it's cause and effect. 00:15:20.323 --> 00:15:32.203 And if if you teach an algorithm cause and effect, then you have to completely negate or move away from the type of learning that we have now in commercial machine learning, commercial machine is purely data driven. 00:15:32.258 --> 00:15:35.768 Now I'm not saying that correlation or data driven doesn't have a purpose. 00:15:35.768 --> 00:15:39.187 I do have a, but it does have a purpose and it would work for different problems. 00:15:39.937 --> 00:15:55.687 for our own, if you want to advance knowledge, as opposed to apply knowledge, is going to be fine for correlation or data driven, because you're looking for a solution you want, see this every day, you want a surrogate model that tells you if you see this, this is likely to happen. 00:15:55.687 --> 00:15:55.988 You know what. 00:15:56.783 --> 00:16:01.043 But if you want to know why things happen, can't rely on AI. 00:16:01.043 --> 00:16:08.753 We have to combine AI into our experiments and we have a completely different kind of teaching methods for AI to figure out cause and effect. 00:16:09.182 --> 00:16:21.452 In one of your papers or in one of your talks you've used in definition of AI as a computational technique that exploits hidden patterns between seemingly unrelated parameters to draw solutions to a given phenomenon. 00:16:21.812 --> 00:16:28.393 But often when you see these data driven AI it just seems like really complex statistics. 00:16:28.393 --> 00:16:35.980 You know, it's like something you could not, plot and having the R square on a single plot is drawn from multiple dimensions, let's say. 00:16:36.519 --> 00:16:39.268 And, This statistical one, it seems attractive. 00:16:39.268 --> 00:16:49.168 It's interesting, possibly useful and probably very useful, but it's this exploitation of hidden patterns between seemingly unrelated parameters. 00:16:49.528 --> 00:17:00.668 This seems like something that could tell us why the facades are burning or why spalling occures, or I don't know why in some conditions, firefighters may die in, in the room. 00:17:00.967 --> 00:17:09.127 So to, uh, but to achieve these hidden pattern uh, recognition, you need knowledge beyond data, right? 00:17:09.127 --> 00:17:11.238 You need to have observations. 00:17:11.238 --> 00:17:13.548 That's what you meant by coupling the experiment and AI AI. 00:17:13.653 --> 00:17:17.403 need to have, you need to have a methodology of saying this. 00:17:18.182 --> 00:17:19.083 have experiments. 00:17:19.083 --> 00:17:23.643 I've seen this, but this experiment is going to be limited by whatever equipment I have sensors. 00:17:23.643 --> 00:17:24.063 They have. 00:17:24.063 --> 00:17:25.885 So sometimes I'm picking up data. 00:17:25.915 --> 00:17:26.786 I think it's noise. 00:17:26.786 --> 00:17:27.776 Maybe it's not noise. 00:17:28.135 --> 00:17:31.286 So you have to do a multiple levels of experimentation. 00:17:32.155 --> 00:17:38.415 Use that data and you have a teacher algorithm at different level what each one means. 00:17:38.415 --> 00:17:45.718 And then the algorithm should be able to put an overall picture of hidden weight pathways between how these factors react. 00:17:46.167 --> 00:17:57.317 For instance, many research papers now on AI and and not just in the fire in really any, any field in engineering, the first, the second user, the second section of a paper but it would be like description of database. 00:17:57.377 --> 00:17:59.278 And then they would list database. 00:17:59.298 --> 00:18:04.278 And then they would say all these that we have min max average median for this, for our database. 00:18:04.817 --> 00:18:15.678 And the second thing that always worked is like a correlation matrix and then they would say, this is the correlation between, the correlation matrix is only going to be linear because you're using a linear correlation. 00:18:15.887 --> 00:18:23.778 There is no guarantee the relationship between the factors themselves or the features themselves or the features and output is linear. 00:18:24.438 --> 00:18:28.644 So having that metrics or that table doesn't really tell me anything about cause and effect. 00:18:28.644 --> 00:18:32.855 It tells me that I could use, any statistical model to come up with an equation. 00:18:33.924 --> 00:18:35.914 the thing about machine learning and statistics is the. 00:18:36.769 --> 00:18:47.059 The statistics are too, as a model you have to be confined with the assumptions of that model because model that has a certain kind of assumptions, applicable for a set of data. 00:18:47.539 --> 00:18:54.380 And hence the statistic can't be applied to many, many of our problems because we have highly on the know problems in machine learning. 00:18:54.380 --> 00:18:59.039 These are non-parametric so they don't have for data. 00:18:59.220 --> 00:19:00.960 We don't have assumption for distribution. 00:19:01.140 --> 00:19:14.055 They can fit very complex functions within our database, most likely they outperform statistics in our case, our problems were linear, you won't find anybody using machine learning because why would you use machine learning Yeah. 00:19:14.140 --> 00:19:15.579 a comfort, such a simple problem. 00:19:16.111 --> 00:19:23.250 I guess in 20 or 30 years, when this becomes like mainstream, you will see people using machine learning for linear problems. 00:19:23.590 --> 00:19:29.240 Just like today, we use a CFD for extremely simple cases instead of zone models. 00:19:29.280 --> 00:19:30.990 I predict is going to happen. 00:19:31.411 --> 00:19:51.799 And, you didn't say that, but I assume that it's also highly related to the amount of data you have to be able to get this quality or this multilevel correlations there, and, like how much that's w when does one know they have enough data for the problem. 00:19:52.309 --> 00:19:58.869 And and another question that's something I would be very interested in when I am planning my experiment. 00:19:59.540 --> 00:20:04.313 How should I prepare myself to create sufficient amount of data, you know? 00:20:04.462 --> 00:20:14.032 So I want to have a grant and I need to know if I will need a thousand experiments, a hundred experiments or 10 experiments of this type and 50 additional of another type. 00:20:14.717 --> 00:20:16.276 Yeah, it's, it's a very good question. 00:20:16.306 --> 00:20:20.476 And it's one, I know there is a lot of research in computer science to figure out this answer. 00:20:20.955 --> 00:20:29.415 know there are a few papers, like I think the 5, 5, 6 years ago, the minimum number of observations somebody would need is maybe I think 10 or 12. 00:20:30.016 --> 00:20:31.605 Now the number is, is 25. 00:20:31.605 --> 00:20:38.836 So if you have a 25 observation, most of the time, you should be able to have some kind of a model that performs in a nice way. 00:20:39.277 --> 00:20:48.506 I know, however, for some type of learning classification problems, think the minimum number that you could be confident about your results is about a hundred observation. 00:20:48.977 --> 00:20:52.369 So these are the numbers that we play with, somewhere between 25 and a hundred. 00:20:52.369 --> 00:21:07.829 Now, the problem becomes is the database becomes wider, you have many, many features, then you will have to have many observation, because if you think about it, a database is a matrix have rows and column, the wider, the matrix so you have to have very, very deep. 00:21:08.349 --> 00:21:11.140 Figured out some kind of correlation between the different parameters. 00:21:11.559 --> 00:21:15.039 So the wider it is the more data that would need. 00:21:15.609 --> 00:21:18.039 And I don't have the, I don't know the answer. 00:21:18.039 --> 00:21:19.930 I don't know if you have 50 data points. 00:21:19.930 --> 00:21:22.269 Is this going to be enough for no, I really don't know. 00:21:22.269 --> 00:21:38.500 I don't think we'll have an answer anytime soon, if we are within, 25 to a hundred points with maybe four to seven features, we should be, or at least the algorithms we have now would be able to give you something that's, maybe with some confidence. 00:21:38.920 --> 00:21:48.279 And to compliment what I just mentioned, , nowadays the algorithms that we do use, they could be augmented with, uh, different tools. 00:21:48.279 --> 00:21:50.940 For instance, you can add confidence intervals to your model. 00:21:51.420 --> 00:21:56.819 So this way, even if you have a small database, algorithm should to the, okay, this is my production. 00:21:56.819 --> 00:22:00.099 I'm predicting this column to fail in 16 minutes. 00:22:01.049 --> 00:22:06.075 this is going to be within a confidence of 90% or 70% of 60%. 00:22:06.404 --> 00:22:12.290 So even if you have a short database, you'd have a prediction, but on the other hand, you have some kind of confidence in your model. 00:22:12.320 --> 00:22:22.040 It's not like the old days, two, three years ago when we couldn't apply confidence intervals were more, and it just, you know, this is the number that you have We could potentially add confidence. 00:22:22.040 --> 00:22:33.740 And then this would give us some level of trust because even if we have a short database or a very, very wide database, that prediction is going to be combined with some kind of confidence that would say, okay, my prediction is this much with 70%. 00:22:34.398 --> 00:22:55.127 In the past, I was, maybe I was not working much with them, but I, I got familiar a bit with some design of experiment methods that allow you to, , figure out the number of experiments to identify the, for example, the influence of variables, on the outcome, Benheken design, , there was aLatin Hypercube, something, Monte-Carlo of course, uh, many, many methods like that. 00:22:55.518 --> 00:23:02.478 And, uh, I was very interested in them because in my PhD, I did, , a very ugly thing. 00:23:02.988 --> 00:23:09.238 I've taken like a hundred geometries, a few fires and some combinations of ventilation. 00:23:09.238 --> 00:23:10.647 And I just brute force to them. 00:23:10.647 --> 00:23:16.498 And it, it gave me a beautiful array of results to work with and complete my PhD. 00:23:16.498 --> 00:23:17.968 And I was very happy with it. 00:23:17.968 --> 00:23:19.917 I'm still, I still am happy with that. 00:23:20.307 --> 00:23:29.394 I just feel like, caveman bashing, a wooden stick against a wall to get an answer where I could have done this way more elegant in a way. 00:23:29.744 --> 00:23:36.404 So, so is there also let's say preparation best practices to drive an experiment. 00:23:36.404 --> 00:23:37.724 So it's useful for this. 00:23:38.387 --> 00:23:59.468 So there are a few things, for instance, we have some algorithms that what all, what they do is they look at the distribution of your observations or the experiments that you have so far, then they are able to zoom or to pinpoint some regions within that distribution that say, okay, this region or for this distribution, we need to have more experimental data points. 00:23:59.468 --> 00:24:01.248 So this way, let's say that Okay. 00:24:01.448 --> 00:24:02.347 to three experiments. 00:24:02.768 --> 00:24:03.518 have to keep in mind. 00:24:03.518 --> 00:24:06.938 Maybe I need to allocate maybe two more experiments with this. 00:24:06.968 --> 00:24:11.567 That would give me To cover a specific, uh, variable for example. 00:24:11.567 --> 00:24:11.928 Okay. 00:24:12.008 --> 00:24:19.738 Because if you think a model and how would validate itself, basically runs some kind of, performance metrics along different regions. 00:24:20.458 --> 00:24:25.008 we normally do is we collect all the data points or the predictions, and we run R square. 00:24:25.587 --> 00:24:26.907 Now R square doesn't tell you. 00:24:27.400 --> 00:24:29.462 the, the performance for specific points. 00:24:29.462 --> 00:24:32.313 It tells you the performance for the whole database or for the whole predictions. 00:24:32.823 --> 00:24:40.502 But if you plot X and Y would see that your curve at some regions are much, much larger than the other regions. 00:24:40.712 --> 00:24:46.202 And for those regions, you might want to end up with experimental points because the distance is two things. 00:24:46.563 --> 00:24:50.883 This one's you want, the algorithm was not able to capture that phenomena at that region. 00:24:50.913 --> 00:24:54.752 And two, it tells you that there is something there that we haven't seen before. 00:24:55.173 --> 00:25:01.653 So maybe if you do an experiment, maybe you'd be able to confirm it or deny it and figure out something new. 00:25:01.726 --> 00:25:08.942 So this will guide you towards the potential outlier that could actually unravel a new physics or something completely unexpected. 00:25:08.942 --> 00:25:09.163 That's. 00:25:09.163 --> 00:25:09.752 That's cool. 00:25:10.272 --> 00:25:13.982 Okay, let's move a bit more into engineering. 00:25:14.712 --> 00:25:33.160 I I've taken a look on your papers and you have used AI to identify fire vulnurable bridges, designed columns, change measurements, for fiber reinforcement, polymers, strengthens reinforced columns to determine spalling it identify failure of beams. 00:25:33.490 --> 00:25:38.289 Is there any field of engineering you have identified that it will not work at all? 00:25:38.559 --> 00:25:39.009 Maybe? 00:25:39.250 --> 00:25:42.099 this is a no, it's a, it's a very good question. 00:25:42.480 --> 00:25:49.720 And then this is what scares me most, to be honest, I, I'm a little bit lucky because I get to play with AI a little bit earlier Yeah. 00:25:49.859 --> 00:25:50.789 to see how it works. 00:25:51.240 --> 00:26:08.303 And so far it's working really well, which tells me it's either the problems we have are enough for AI, because if you really think about it, scientists, when they develop an algorithm it works for insurance, for medicine, for space, it's not just, know, bending. 00:26:08.333 --> 00:26:10.313 Buckling, flammability, collapses. 00:26:10.792 --> 00:26:12.323 have much, much more complex problems. 00:26:12.833 --> 00:26:19.086 if it works well for complex problems, like finding a new star galaxies, that's a very complex problem. 00:26:19.566 --> 00:26:29.893 Maybe our problems not that after all, for AI to solve, maybe the, maybe it's complex for empirical methods that we apply or for finite elements methods that we use. 00:26:30.002 --> 00:26:31.063 They're not really that hard. 00:26:31.063 --> 00:26:34.093 They're just computationally expensive for FEA or like FDS. 00:26:34.153 --> 00:26:35.682 You have to run them for a long time. 00:26:35.682 --> 00:26:41.442 You have to mention it had elements, but collectively, maybe they're not the other issue I'm thinking. 00:26:41.442 --> 00:26:48.455 I mean, I think about is maybe because algorithm is really a black box and we don't know why, how it behaves. 00:26:48.873 --> 00:26:49.682 We only see the. 00:26:50.522 --> 00:26:56.252 What if the output is correct, but the map that links the inputs to the output is not correct. 00:26:56.762 --> 00:27:04.442 Maybe the output is correct for this database, but once you, once you go for outside of the range of database, it's going to be very, very hard. 00:27:04.472 --> 00:27:09.182 Or maybe you start to get some errors that the counterpart is the following. 00:27:09.833 --> 00:27:16.613 The counterpart is very interesting because let's say in structural fire engineering, like we have certain sizes for columns and beams. 00:27:17.103 --> 00:27:18.123 don't go beyond that. 00:27:18.807 --> 00:27:19.127 Hm. 00:27:19.173 --> 00:27:24.663 databases are usually good because, you won't find a very, very, very thin column or a very, very short column. 00:27:24.663 --> 00:27:25.413 We don't use that. 00:27:26.042 --> 00:27:31.836 going outside the norm will give you error measurements, but at the same time, we never used in practice. 00:27:32.496 --> 00:27:40.654 there is like a and cons for, for every issue In the same way, uh, you will have only combustion within limits of flammability. 00:27:40.924 --> 00:27:45.454 You would have certain sizes of fire only in certain, uh, ventilation factors. 00:27:45.785 --> 00:27:51.724 So there are like boundaries to the fires that we know empirically and we could work with that. 00:27:51.922 --> 00:28:03.288 With all these innovation that you show in machine learning, I'm really wondering how hard is it in a field so let's say a field with concrete, like ours. 00:28:03.528 --> 00:28:04.877 Like, uh, construction. 00:28:05.057 --> 00:28:07.877 It's not a place of raging innovation. 00:28:07.968 --> 00:28:15.471 It's, we're using hundred years old standards to quantify a fire resistance and it's not unlikely to change very soon. 00:28:15.891 --> 00:28:18.681 The problem is not with innovating something. 00:28:18.891 --> 00:28:29.171 The problem is innovating something and not breaking, everything else, and we're very slow to adopt new technologies, new methods, new. 00:28:29.330 --> 00:28:34.010 How is it going for you as a pioneer of this technology in construction? 00:28:34.317 --> 00:28:36.627 You must have a funny reviews for, for your papers. 00:28:37.835 --> 00:28:42.275 Um, papers, I would get very unique reviews. 00:28:42.275 --> 00:28:42.724 Yes. 00:28:43.085 --> 00:28:58.085 I would say now it's much better now and it much, much better, I would say the following a hundred percent with a slow to adapt, I would agree with this, like two, two years ago, when I was trying to push for something in a conference or with a funding agency, the answer is not complete. 00:28:58.144 --> 00:29:03.244 So I had to reconsider my whole path because I couldn't get anything from anybody. 00:29:03.875 --> 00:29:06.434 And nowadays I think the industry is interest. 00:29:07.123 --> 00:29:10.685 Like we, we had, with the American concrete Institute, we had a few talks. 00:29:10.685 --> 00:29:13.685 We have, we published a book with them on AI, completely with concrete. 00:29:13.685 --> 00:29:15.125 They were very, very open to it. 00:29:15.201 --> 00:29:15.671 okay. 00:29:15.806 --> 00:29:24.486 the steel industry is looking to something in a close by mass and we actually, one of my grants is from masonry they want to use AI to design masonry structure for fire. 00:29:25.415 --> 00:29:30.185 it's a little bit bad, but that the thing that's good for our case right now is the following. 00:29:30.546 --> 00:29:33.349 We have a lot of startups many of these are startups. 00:29:33.380 --> 00:29:38.869 They're trying to automate many of the routine applications or that we use. 00:29:39.440 --> 00:29:46.009 then these are a way to automate our routine step is to use machine learning because you know, you're not really going above and beyond in something. 00:29:46.339 --> 00:29:48.170 You're you have a procedure already. 00:29:48.170 --> 00:29:54.920 You're just trying to make it much, much faster, much more accurate, less error, all of that would accumulate to less time, more money. 00:29:55.700 --> 00:30:01.460 So there is a push from the industry now, and I think it will grow within the next two or three years because there's a lot of startups. 00:30:01.970 --> 00:30:02.750 these are startups. 00:30:02.779 --> 00:30:04.099 If you look at the investors with. 00:30:05.009 --> 00:30:06.230 They're not really engineers. 00:30:06.289 --> 00:30:10.339 They're mainly into developing softwares and apps and computer science backgrounds. 00:30:10.940 --> 00:30:13.940 however, they don't have the domain knowledge that we do. 00:30:13.940 --> 00:30:17.930 And this is why they hire civil engineers or structural engineers or fire engineers. 00:30:18.680 --> 00:30:22.202 they learn the problem, it's going to be very, very easy for them to develop solutions. 00:30:22.232 --> 00:30:24.752 The issue with me is the following. 00:30:24.752 --> 00:30:32.836 The issue is we need solutions that come from somebody who has been practicing and been educated in our field solve our problems. 00:30:33.316 --> 00:30:40.185 We can just give the domain knowledge to somebody who's doesn't have our background because they're looking at the surface. 00:30:40.185 --> 00:30:44.553 We need people to do a fire engineering or structural engineering from the beginning, from the under. 00:30:45.288 --> 00:30:51.887 Then you would come up with solution that would work better, would work best for our case and will advance our knowledge as well. 00:30:51.917 --> 00:30:56.657 Because you know, I'm not really looking for a software that tells me is the amount of fire you're going to get. 00:30:56.657 --> 00:31:00.567 Or like, this is the heat intensity you're going to get, because anybody can do a software like this. 00:31:00.807 --> 00:31:04.857 I want to know why, if I know why redesign, I can change things. 00:31:04.857 --> 00:31:09.087 I can come up with unique designs, innovations that we don't have right now. 00:31:09.478 --> 00:31:12.928 And the computer scientists can give you that you have to be an engineer of that. 00:31:13.444 --> 00:31:20.065 I'll challenge that because,, for a paradigm shift to occure in a field, it must be done by someone from outside, outside of that field. 00:31:20.144 --> 00:31:28.194 If you're a graduated fire safety engineering, it's very unlikely that you will change the fire safety engineering completely because of the way how you would have been thought. 00:31:28.194 --> 00:31:36.414 And there is certain experience factor in your computer in your head that, uh, will prevent you from touching stuff. 00:31:36.444 --> 00:31:47.934 But our, you escaping the field and jumping into computer science and coming back is actually quite a nice path to, to carve such a path for for new, uh, if also you stern black books many times. 00:31:47.934 --> 00:31:50.318 And, it seems like that, I don't understand it. 00:31:50.318 --> 00:31:54.459 You know, I see, I know I can put some stuff in, it will give me stuff out. 00:31:54.701 --> 00:31:55.781 Even for CFD. 00:31:55.781 --> 00:32:01.011 I'm in CFDs very, very hard, but I can more or less understand CFD. 00:32:01.031 --> 00:32:01.602 I don't claim. 00:32:01.602 --> 00:32:06.912 I understand it completely, but I, more or less know what the equations do, what the schemes are. 00:32:06.922 --> 00:32:09.912 What's a turbulence model, what's boundary layer. 00:32:09.942 --> 00:32:15.672 I know these things and I can track back my simulation, identifying each of these steps and going in. 00:32:16.001 --> 00:32:19.842 And then I see a pattern of neural network and it looks like a Christmas tree to me. 00:32:20.082 --> 00:32:23.412 It doesn't reassemble a anything equation or something. 00:32:23.412 --> 00:32:33.082 And, in your paper, in, um, Automation in Construction, engineers guide to AI, you you've championed this explainable AI as a necessity. 00:32:33.692 --> 00:32:41.461 So tell me what would be this explainable AI and why it would be something that would make me use AI and while I'm not using it today. 00:32:41.852 --> 00:32:42.031 Yeah. 00:32:42.632 --> 00:32:45.122 So what I did is it's, it's a very simple exercise. 00:32:45.122 --> 00:32:54.842 So you take an equation from a code and you apply in a database and you see that the equation from the code that we have to use as engineers does not perform as well as an algorithm. 00:32:55.321 --> 00:32:58.082 So this fact by it sort of should make you pause. 00:32:58.082 --> 00:33:01.471 How can you not trust a code over an algorithm? 00:33:02.281 --> 00:33:13.204 Then the second question would be if the algorithm can predict better than the code, why do I have to use the code when they have a better method that can predict better than the code? 00:33:13.805 --> 00:33:15.541 The second question is the following. 00:33:16.172 --> 00:33:18.541 Why does the algorithm do that? 00:33:19.051 --> 00:33:19.652 the code cannot. 00:33:20.691 --> 00:33:23.277 Now to know why an algorithm does a certain thing. 00:33:23.277 --> 00:33:27.567 We have to break that, and see how it, how it does the way it does. 00:33:28.017 --> 00:33:34.136 And right now we don't know we can do that because one would not commit a scientist even computer scientists. 00:33:34.166 --> 00:33:44.517 Can't really track how the algorithms work, because everything for them is goal is to get as good of a prediction as an experiment, or as an observation, don't care how to get there. 00:33:44.517 --> 00:33:47.636 In our case, we do care because we have to justify our decisions. 00:33:47.936 --> 00:33:52.977 How can they justify using a column with two hour fire rating in the building they don't know why? 00:33:52.977 --> 00:34:03.297 If the algorithm says, yes, I, if I know why need to figure out what, so this is when, do you have to use If you have AI, you have also above it, explainable AI, explainable AI. 00:34:03.866 --> 00:34:05.096 The way it does is the following. 00:34:05.186 --> 00:34:09.896 Each algorithm should be able to tell you exactly how it came up with its own production. 00:34:10.527 --> 00:34:11.996 It has to break it down for you. 00:34:11.996 --> 00:34:15.699 So you can understand because numerically It's correct. 00:34:15.699 --> 00:34:19.960 However, physically, or from an engineering maybe it's not correct. 00:34:20.679 --> 00:34:27.494 the algorithm of say the relationship between, you know, material and geometry is, is linear, but we know from our experiment, it's not linear. 00:34:27.943 --> 00:34:29.293 how can I trust it to production? 00:34:29.344 --> 00:34:31.384 If it negates what physics tells me to do. 00:34:32.074 --> 00:34:34.534 Now, the problem with explainable AI is the following. 00:34:34.976 --> 00:34:39.056 It will only explain its results based on the database that you have. 00:34:39.117 --> 00:34:53.067 So if you don't have a good database or as many features as the physics would allow you to do, even if you have explainable AI, it won't be as good as the one we have in physics, because it's, won't be able to capture all the interactions that one we see in physics. 00:34:53.367 --> 00:34:56.246 So it's not really about using explainable AI or AI. 00:34:56.246 --> 00:34:59.786 It's about using a system that can tell you this because. 00:35:00.432 --> 00:35:07.956 When we use explainable AI, we basically have a very small code within our algorithm that can track prediction back to its origin. 00:35:08.527 --> 00:35:14.893 did the algorithm link, parameter one with parameter two with parameter three with parameter, for to come up with a prediction that it did. 00:35:15.117 --> 00:35:15.956 Black box. 00:35:15.987 --> 00:35:41.693 Doesn't tell you that So for example, if I employed, AI to predict, , smoke movement in a, let's say buoyant plume, it could actually, in the meantime, tell me that it works when you assume the gravity is less on the Mars and then it works while in fact, it's just a matter of entrainment coefficient that is elsewhere with which could accidentally be the same number as the ratio of gravity here in the Mars. 00:35:42.023 --> 00:35:44.934 But then the algorithm will never know what happened. 00:35:44.934 --> 00:35:46.824 It just used this and it worked. 00:35:46.824 --> 00:35:49.360 And for them, it's, perfect for engineering. 00:35:49.360 --> 00:35:58.260 You need to understand, and this, uh, so breaking it into steps and seeing the more or less what has been done gives you this. 00:35:58.289 --> 00:36:05.010 Let's say higher power to unravel this hidden patterns and you care less about advanced statistics. 00:36:05.099 --> 00:36:05.550 exactly. 00:36:05.550 --> 00:36:16.650 And the other thing is, at least in my eyes, if I know how the algorithm sees the problem, I might be able to figure out any phenomena or sub phenomena that they haven't known before. 00:36:16.650 --> 00:36:18.510 And maybe this is why I'm very can methods. 00:36:18.666 --> 00:36:25.777 They're by design, very conservative because we have to be conservative, maybe we could, if we know why we don't have to be extremely conservative. 00:36:25.867 --> 00:36:38.166 And plus we know now something new that we didn't know before we can figure out why this thing, when I think of AI I always think of, I tool that can give me an answer to why this thing I did to why I didn't know this before. 00:36:38.887 --> 00:36:40.356 What's this new knowledge to me. 00:36:40.356 --> 00:36:42.483 I'm not really looking for c orrelation. 00:36:42.516 --> 00:36:58.606 I mean, my earlier work was heavily data-driven correlation because I mean, I didn't know better, but nowadays it's just makes more sense for some papers, of course data-driven would work because paper itself is for a data driven problem, the overall idea should not always be data driven. 00:36:58.606 --> 00:37:00.407 It should be more, much more than that. 00:37:00.436 --> 00:37:02.327 It should always be advanced in science. 00:37:02.356 --> 00:37:07.516 How can we advance our knowledge having to spend and thousands of dollars? 00:37:07.516 --> 00:37:12.277 And, and here's the thing, Somebody that does experiment now, years from now, it's forgotten. 00:37:12.527 --> 00:37:16.836 Somebody goes back and repeat the experiment, and they get the grant to redo the experiment again. 00:37:16.867 --> 00:37:18.637 Or they don't, you know, expand an experiment. 00:37:19.266 --> 00:37:23.617 papers wouldn't be published something in a way after a few months, it's shelved away. 00:37:23.786 --> 00:37:28.327 It's in a database sciencedirect or Springer, it's, it's being online. 00:37:28.447 --> 00:37:29.496 We rarely visit. 00:37:30.150 --> 00:37:32.820 But why do we have to continue doing the cycle all over again? 00:37:32.820 --> 00:37:39.539 If you accumulate our knowledge and we are able to come up with something new, then we can different directions that we haven't seen before. 00:37:40.384 --> 00:38:00.278 I've started with classifying this, AI into supervised unsupervised semi-supervised and you were talking about regression classification on other ways to formally classify this, but I think the true first choice is, do you use it for discovery or you do it to calculate something, you know? 00:38:00.309 --> 00:38:15.376 And, I think that's the first thing, because if you just want to figure out a number out of a very complex array of results, you have obtained that you're unable to process, in other way, because the correlations or somethingare multidimmensional. 00:38:15.721 --> 00:38:25.818 then you probably are seeking a different path than when you try to employ this method to do, to find unexpected and discover something. 00:38:25.951 --> 00:38:29.451 and, as an engineer, I would like to have better numbers. 00:38:29.451 --> 00:38:41.181 I would not necessarily be happy, discovering a completely new failure mode because that's, uh, I mean, I made the wrong, or we're kind of screwed as a humanity if I do. 00:38:41.670 --> 00:38:48.221 But as a scientist, I would like, I maybe care less about the numbers and they would care more about discovery. 00:38:48.641 --> 00:38:52.643 And, coming back to your thought about,, collectively adding to that. 00:38:53.393 --> 00:38:53.664 Okay. 00:38:53.713 --> 00:38:55.844 To what extent the data from the past. 00:38:56.398 --> 00:38:56.938 Exactly. 00:38:57.193 --> 00:39:00.943 To what extent you can take your papers from 50, 60 seventies. 00:39:00.974 --> 00:39:08.353 I don't know, from last IAFSS and use them to develop your own models is how big of an issue is that? 00:39:08.483 --> 00:39:12.923 The thing is because we talk about fire, it's, a very read it. 00:39:12.954 --> 00:39:16.164 Like it's a very niche area that, and it's a very expensive area. 00:39:16.164 --> 00:39:31.797 We don't really have a lot of experiments, or like low-risk tests that we can use, but we do have some, if you want to start with machine learning with a goal to come up, let's say with a black box surrogate that tells you failure mode or failure time, rather than doing a very lengthy calculation. 00:39:33.027 --> 00:39:39.387 You really have to do what you have, and those would be experiments that the old experiments now, the good thing is the following. 00:39:39.867 --> 00:39:45.141 The good thing is those experiments are the same ones that we use now by our care. 00:39:45.141 --> 00:39:51.291 So, you know, in a way, w we have some kind of similarity, the, on the other, on the opposite side material is different. 00:39:51.291 --> 00:39:55.130 Like for instance, concrete 50 years ago is, is really different than the concrete we have now. 00:39:55.161 --> 00:40:02.121 So the experiments 50 years ago which is also, which is what the codes are built on, are not built on your new experiments. 00:40:02.686 --> 00:40:03.266 I'm happy. 00:40:03.266 --> 00:40:03.985 You've added that. 00:40:04.201 --> 00:40:07.771 Codes are built on very, very old expert in the sixties, fifties, seventies. 00:40:08.010 --> 00:40:14.641 So even the code, why would you apply the code now when it's 60 years later, how does that accumulate to what we have now? 00:40:15.096 --> 00:40:15.456 Wow. 00:40:15.661 --> 00:40:22.068 a way, I know it may not be as comprehensive or as accurate as doing the knowledge that we have now. 00:40:22.398 --> 00:40:40.596 However, this is the practice that we're using And if you want to compare, if you really want to compare, let's say code that a procedure against machine learning to be fair, you have to kind of use the data that develop that that a provision, which is the old data and apply to the algorithm and see, the comparison. 00:40:40.596 --> 00:40:44.666 You have to have a fair line of comparison to, this is how it works but however,. 00:40:44.735 --> 00:40:47.976 Am I happy with using 50 year old experiments? 00:40:48.096 --> 00:40:49.025 I'm not happy. 00:40:49.025 --> 00:40:50.465 No, but this is the ones we have. 00:40:50.465 --> 00:40:51.936 And this is the standard we have to use. 00:40:52.175 --> 00:40:59.088 Maybe in the future, it will be different on the good side, on the other dimension, using all the experiments. 00:40:59.213 --> 00:41:02.483 And let's say that you have two columns, one very, very old one, very, very new. 00:41:03.076 --> 00:41:05.759 The failure mode is not going to be something new. 00:41:06.509 --> 00:41:08.398 It would have still failed in the same manner. 00:41:08.789 --> 00:41:12.688 However, the time Um, it's going to be different because we have different chemicals. 00:41:12.688 --> 00:41:15.088 We have different stuff now that we use in our material. 00:41:15.266 --> 00:41:17.358 at different loading, different, temperature. 00:41:17.699 --> 00:41:21.748 However, the we're subjecting this element to is the same. 00:41:21.748 --> 00:41:23.619 They still have the same chamber compartment. 00:41:24.159 --> 00:41:27.398 So it's not really that we're completely using something different. 00:41:27.878 --> 00:41:29.469 It's just, there are some differences. 00:41:29.469 --> 00:41:35.559 And even if you want to do a, like a statistical analysis, like a meta analysis, have to compare different data from different experiments. 00:41:35.739 --> 00:41:53.646 And this is, again, this is why using data-driven analysis is a little bit itchy for me now, because I want to know why, like, at least in my mind, I want to know why this column fails so if it's fail that, years ago, or now there has the mechanism is not going to be some, some new physics. 00:41:53.746 --> 00:41:56.315 It's going to be something that maybe we haven't seen before. 00:41:56.916 --> 00:41:57.965 how can I get to that? 00:41:57.965 --> 00:42:18.427 Something that we haven't seen before by using the same old methods that we have been using for 50 years by now, we would have figured out, you know, So maybe if we use something, method, maybe we can see a little bit different and maybe that a little bit of difference would open up a new experiments for us or a new research area for us that we can apply news That's interesting. 00:42:18.427 --> 00:42:23.340 And, for, for me personally, I really liked the Xinyans, I'm in the world of smoke control. 00:42:23.340 --> 00:42:39.574 And I really loved, how he perceived that the CFD could let to let's say more capable algorithms that would predict, the smoke behavior in a compartment, giving you a number the time to, for the layer to fall down or some tenability criteria to be breached. 00:42:40.143 --> 00:42:56.547 And for example, one of mine main, , areas of research is car parks . I engineer a lot of smoke control in car parks and our limitations is usually that we take a car park, we do 2, 3, 4, 5 CDs in it, for a certain size of the fire. 00:42:57.148 --> 00:43:01.108 And I assume if I did a sufficiently large amount of. 00:43:01.572 --> 00:43:02.552 Simulations. 00:43:02.913 --> 00:43:06.242 And then they have received a new car park with the new architecture. 00:43:06.242 --> 00:43:11.822 With, I know I've performed 1, 2, 3 simulations in that carpark. 00:43:12.182 --> 00:43:18.813 The algorithm could technically take over and tell me what would happen in like a thousand different scenarios in that car park. 00:43:18.992 --> 00:43:20.905 could you use it in like, this, like. 00:43:21.085 --> 00:43:21.206 Yeah. 00:43:21.775 --> 00:43:27.186 instance, now what we're doing, we have a database of our, say columns, 200 columns. 00:43:27.266 --> 00:43:33.452 We could come up, we can ask the algorithm to simulate a data worth of testing, 5,000 columns or 10,000. 00:43:34.338 --> 00:43:41.148 So this way, I'm trying to capture as many interactions between the features or between the parameters as I couldn't have done using experiments. 00:43:41.748 --> 00:43:47.688 But however, I still to have that baseline that at least as the algorithm, this is the main, this is the map. 00:43:47.728 --> 00:43:52.847 This is the average of the distribution of the possible that I could see before. 00:43:53.202 --> 00:43:53.532 Hm. 00:43:53.557 --> 00:43:53.927 Xinyan. 00:43:54.311 --> 00:44:02.617 Any the problem that you're going to be using simulation for it will be expensive not only expensive, you'll have to continue to do it all over and over again. 00:44:03.307 --> 00:44:09.623 And you know, once we're done with this, say with your design, you throw away this, you know, maybe you clear your desks and yeah. 00:44:09.623 --> 00:44:09.882 Yeah. 00:44:09.887 --> 00:44:10.367 throw it away. 00:44:10.547 --> 00:44:37.643 But if you have this, let's say on an annual basis, and let's say you designed 50 structures or like 50 cases, the simulation that you have is very valuable information because if you accumulate them by five or six or seven years, you'll have a very, very good database that you can teach an algorithm Maybe figured out something that we haven't seen before, maybe come up with some kind of a faster approach to solve the small problem to figure out at least what could be. 00:44:37.673 --> 00:44:42.233 And this would be interesting to me, what would be a severe case for this parking structure? 00:44:42.893 --> 00:44:51.413 it without having to house it or doing many minutes in a CFDs before, I may be wrong, but do you exactly know which one would be a severe case right now of hand? 00:44:51.463 --> 00:44:54.032 It is expert judgment and you use the design. 00:44:54.932 --> 00:45:04.503 That's the thing, because if you could use this technique to expand the number of investigated cases, you can start talking about risk and probabilities. 00:45:05.163 --> 00:45:10.742 Like a fire of this probability is giving these consequences with this confidence. 00:45:11.132 --> 00:45:16.885 And the fire of this probability is giving you these consequences at these intervals, and then you go, and it was beautiful. 00:45:16.885 --> 00:45:24.585 You could ask the AI, please test any smoke exhaust capacity from this amount of CFMs to this amount of CFMs. 00:45:24.976 --> 00:45:32.635 And, then it will tell you, okay, if you increase the ventilation twice, you decrease your probabilities by this amount. 00:45:32.635 --> 00:45:36.985 And if you increase it sevenfold, it doesn't change much from the previous case. 00:45:37.465 --> 00:45:43.076 So you, you start to get much more detailed. 00:45:43.230 --> 00:45:53.460 Outcome of your analysis then you would have from investigating multiple points, even if you are the best CFD engineer in the world, because it's not the tool that limits you. 00:45:53.880 --> 00:46:02.760 It's the capabilities of running multiple parallel cases that's essentially limiting and that there was a solution to that. 00:46:02.789 --> 00:46:04.170 There exists solutions to that. 00:46:04.590 --> 00:46:07.907 There was PhD student of Bart Merci and now Dr. 00:46:08.056 --> 00:46:13.166 Bart van Weyenberge, who was doing his PhD on a response surface technique. 00:46:13.467 --> 00:46:21.713 It's a statistical technique where you can map, certain, uh, inputs to certain outputs of, multidimensional, uh, surfaces. 00:46:22.342 --> 00:46:24.233 And from that you can buy running. 00:46:24.713 --> 00:46:35.632 Let's say 10,CFDs or 20 CFDs you can predict the outcome of multiple CFDs, but it still requires you to solve for a certain geometry in here with machine learning. 00:46:35.632 --> 00:46:42.543 Maybe you could use results from different building to enhance your knowledge about this particular building. 00:46:42.559 --> 00:46:45.099 I mean, it's amazing because it already did. 00:46:45.099 --> 00:46:49.179 The response surface seemed like magic, and this is magic plus. 00:46:49.480 --> 00:46:52.380 It's if that happens it's gonna be amazing. 00:46:52.650 --> 00:46:53.880 And I really wish It happened. 00:46:54.869 --> 00:46:58.612 So if I wanted it, to happen What should I do now? 00:46:58.822 --> 00:47:00.081 Should I go learn coding? 00:47:00.101 --> 00:47:01.121 what's the first step? 00:47:01.152 --> 00:47:01.947 And, Let's assume. 00:47:01.947 --> 00:47:03.597 I don't know anything about coding. 00:47:03.597 --> 00:47:03.987 I don't know. 00:47:03.987 --> 00:47:04.378 Python. 00:47:04.378 --> 00:47:04.768 I don't know. 00:47:04.768 --> 00:47:06.018 R I don't know anything. 00:47:06.315 --> 00:47:07.402 but I just love this. 00:47:07.431 --> 00:47:09.742 Where, where should they, what should they do with myself? 00:47:09.891 --> 00:47:11.842 to me honest, I learned everything on YouTube. 00:47:12.081 --> 00:47:13.311 They have five Yeah. 00:47:13.791 --> 00:47:15.952 for every kind of For everything. 00:47:15.981 --> 00:47:16.192 Yeah. 00:47:16.702 --> 00:47:17.661 can either learn from them. 00:47:17.692 --> 00:47:21.891 The good news is Uh, at this moment, we don't really develop algorithms. 00:47:21.891 --> 00:47:23.632 So there are many, many codes. 00:47:23.641 --> 00:47:26.842 Like if you go to SciKit there is like the cause already there. 00:47:26.842 --> 00:47:28.831 So you can just copy paste them, Hmm. 00:47:28.851 --> 00:47:29.992 add your data, run it. 00:47:30.001 --> 00:47:32.864 And then, you know, you can fine tune, a few parameters. 00:47:32.925 --> 00:47:34.005 You should be good to go. 00:47:34.005 --> 00:47:35.594 It's not, it's something that's complex. 00:47:35.925 --> 00:47:50.954 Once you start to go maybe into explainability confidence, trust, then you have to have some kind of a very good background when it comes to math or calculus, because they're, at that point, it's not just ask them, but as applying, in other words, it's more on the development side. 00:47:51.585 --> 00:47:56.324 if you want to figure it out causality, or for instance, cause and effect, this is at least what I'm trying to do. 00:47:56.715 --> 00:48:01.414 you want to figure out cause and effect, then you have to have much more higher advancements for coding. 00:48:01.894 --> 00:48:04.125 So the bottom box, I mean, this is what i do with my students. 00:48:04.980 --> 00:48:08.550 I'm not really expecting you to a new algorithm. 00:48:08.550 --> 00:48:09.869 If we can do that, that'd be great. 00:48:10.260 --> 00:48:14.340 However, the algorithms we have now can solve many, many, many problems. 00:48:14.369 --> 00:48:16.500 And all what you really have to do is two things. 00:48:16.500 --> 00:48:22.079 One understand how the algorithm worked its assumptions as limitation know how to apply it. 00:48:22.728 --> 00:48:29.099 You don't read it need to code it by hand because the codes are already they're available online. 00:48:29.099 --> 00:48:30.570 You can just copy paste them from there. 00:48:31.230 --> 00:48:32.760 to find your data and apply it. 00:48:32.760 --> 00:48:34.920 And then will see if you apply it. 00:48:35.340 --> 00:48:37.559 mean, I did this experiment in two of my favorites. 00:48:38.159 --> 00:48:41.820 took five or six algorithms and I applied them by default values. 00:48:41.820 --> 00:48:45.179 I just copied and pasted them on our data. 00:48:45.579 --> 00:48:46.400 And it works. 00:48:46.530 --> 00:48:59.378 I mean, you get 95%, you get 90% with very, very cheap resources, tells me that you could basically apply the same algorithm for different problems and your are gonna get very good results too. 00:48:59.831 --> 00:49:04.974 Not all the time, but at least for the most of the time, because these algorithms are extremely powerful. 00:49:05.074 --> 00:49:15.454 That's a relief in a way, you know, and I had Matt Bonner as well in here and he told the same thing that there are algorithms that exist and you can apply them. 00:49:16.054 --> 00:49:17.135 Xinyan said the same. 00:49:17.135 --> 00:49:18.925 You're the , third person to tell you the same. 00:49:18.925 --> 00:49:24.574 So I must build my brave and, and just try, I guess that's how I learned programming. 00:49:24.605 --> 00:49:29.074 Actually just, just keep trying and do as many mistakes as you can. 00:49:29.074 --> 00:49:30.724 And eventually it will work out. 00:49:31.235 --> 00:49:31.534 So. 00:49:31.670 --> 00:49:34.130 a new course next fall on machine learning. 00:49:34.849 --> 00:49:38.255 send you a link for my lecture so you can, you can attend Oh, really? 00:49:39.461 --> 00:49:40.302 That's so cool. 00:49:40.541 --> 00:49:41.742 I would appreciate that. 00:49:42.242 --> 00:49:48.032 for the end that you usually referring to resources and you have your webpage, that's very rich in resources. 00:49:48.032 --> 00:49:50.501 So I will also link to that. 00:49:50.532 --> 00:49:56.405 And, you had the paper in Fire Technology about, different types of machine learning that can be used in fire. 00:49:56.764 --> 00:50:02.885 You had this Engineer's Guide to AI and automation in construction, which was a very interesting case study. 00:50:02.885 --> 00:50:04.965 And, , it was a really nice paper. 00:50:04.994 --> 00:50:16.105 W what else should I refer the audience to, to read up on, on this, I really feel , for fire there is going to be, the mechanistic, , review paper is a very good one , for a beginner. 00:50:16.164 --> 00:50:20.244 I know I sent out professor Rein in like a very short letter. 00:50:20.485 --> 00:50:21.925 It's going to be published very, very soon. 00:50:21.925 --> 00:50:23.875 So that, that would be a compliment to that one. 00:50:23.875 --> 00:50:25.795 Once it is, I'll send you a link for that. 00:50:26.581 --> 00:50:33.481 Engineer's Guide is, is one of my PIs, or at least when I think about it, this was highlight of 2021. 00:50:33.481 --> 00:50:35.711 For me, that, that paper Is the one really? 00:50:36.641 --> 00:50:37.452 I really liked it. 00:50:37.452 --> 00:50:37.722 There. 00:50:37.952 --> 00:50:43.621 it, even the times that I spent a lot of time on the title, because I figured that would be something very close to my heart. 00:50:44.041 --> 00:50:46.742 Uh, there is a third paper it would be mapping function. 00:50:46.742 --> 00:50:49.262 So it, I think it's after naming function. 00:50:49.831 --> 00:50:58.172 this is where we're trying to use more of a cause and effect kind of machine learning, or how can we arrive at that cause and effect having to hassle with coding. 00:50:58.172 --> 00:51:15.172 And we can actually figure out a pathways between different algorithms come up with a function or mathematical expression that can convey to us some kind of, a formula or at least can give us, because if you think of the, out of the fanatical with them, it's a number that for us engineers would like to see. 00:51:15.291 --> 00:51:16.632 We're trained in formulas. 00:51:16.751 --> 00:51:25.105 We see that for instance, this is the format that you can apply get an output machine and it gives you a number and hence, this is the hesitation. 00:51:25.105 --> 00:51:27.565 We can see why we can see how it was. 00:51:28.092 --> 00:51:29.141 In mapping function. 00:51:29.141 --> 00:51:35.952 It's a way that it can translate algorithmic logic from a black box into a function that we can see. 00:51:36.431 --> 00:51:46.628 if you can see it, you can see the interaction between the parameters, your life had to feel much more comfortable applying a function, as opposed to applying a complete black box that we don't know why does okay. 00:51:46.628 --> 00:51:48.128 That's really, really good. 00:51:48.128 --> 00:51:53.768 And, some external or, uh, resources like maybe YouTube channel or something that you can recommend send you these. 00:51:53.768 --> 00:51:54.998 I have them on my bookmarks. 00:51:54.998 --> 00:51:58.400 I'll send you a, really you a link Fantastic. 00:51:58.400 --> 00:52:02.932 I'll put it in the show notes and I, I hope, , someone will, find it useful. 00:52:02.932 --> 00:52:07.882 And, uh, I really, I really appreciate you, you sharing this knowledge okay. 00:52:07.882 --> 00:52:09.932 Nasser, that was a great talk. 00:52:09.932 --> 00:52:18.722 And I learned something about AI today and, maybe I'm one step closer to understanding how it can be applied in my field. 00:52:18.722 --> 00:52:24.485 And I guess there's many heads buzzing now, how can this be implemented in their fields? 00:52:24.940 --> 00:52:25.360 okay. 00:52:25.481 --> 00:52:29.320 Thank you for joining us in the Fire Science Show and I hope you had a great time. 00:52:29.320 --> 00:52:31.070 I had the lot, Thank you very much. 00:52:31.070 --> 00:52:32.360 I appreciate your reaching out. 00:52:32.481 --> 00:52:33.351 I appreciate your show. 00:52:33.351 --> 00:52:33.771 Very good. 00:52:33.771 --> 00:52:36.490 I mean, I always watch watch the shows when you post them on Twitter. 00:52:36.521 --> 00:52:38.101 It's a very really? 00:52:38.190 --> 00:52:38.820 That's cool. 00:52:39.501 --> 00:52:40.130 I like that. 00:52:40.130 --> 00:52:42.291 You just do one thing. 00:52:42.380 --> 00:52:46.190 It's like different components within the fire wrodl so it's much more informative this way. 00:52:46.561 --> 00:52:46.860 Yeah. 00:52:47.121 --> 00:52:47.840 that very much. 00:52:48.036 --> 00:52:48.817 Thank you so much. 00:52:48.847 --> 00:52:49.356 Cheers, man. 00:52:49.436 --> 00:52:49.706 Bye-bye. 00:52:50.802 --> 00:52:51.432 And that's it. 00:52:51.603 --> 00:52:53.233 Well, what a discussion that was. 00:52:53.282 --> 00:52:58.193 Maybe I just should open some python right now and then start digging into that. 00:52:58.193 --> 00:53:03.262 I'm really excited about this world of fire science and the, possibilities it brings. 00:53:03.922 --> 00:53:18.378 MZ has used AI in so many different aspects of fire engineering, like literally go to his webpage and check out his papers, the variety of topics, where this method was used and considered useful. 00:53:18.643 --> 00:53:22.063 It's just amazing how wide this technology is. 00:53:22.422 --> 00:53:23.563 Of course there are caveats. 00:53:23.893 --> 00:53:26.563 You need to worry about the data quality. 00:53:26.563 --> 00:53:29.083 You need to worry about what the algorithms have not seen. 00:53:29.413 --> 00:53:32.023 I hope you've picked up these things from our discussion. 00:53:32.023 --> 00:53:37.652 That technology is powerful, but just as powerful as the algorithm and as powerful as the data that fuels it. 00:53:38.163 --> 00:53:42.702 And by far, most importantly, as powerful as the person who's using that. 00:53:43.123 --> 00:53:43.873 So if. 00:53:44.733 --> 00:53:47.682 I don't know what you're doing and you drop machine learning on that. 00:53:48.132 --> 00:53:58.097 Well, you're going to have a machine learned no idea what you're doing, but if you know what you are doing and you know what you're looking for, it's just hell, a complex problem to dig into that. 00:53:58.516 --> 00:54:04.086 Well then machine learning and artificial intelligence, maybe your best future friend. 00:54:04.679 --> 00:54:09.635 Now this talk today, I think it's a part of a mini series in the podcast. 00:54:10.175 --> 00:54:24.034 If you remember, I had an episode with Xinyan Huang from Hong Kong Polytechnic university, with whom I have discussed artificial intelligence and its potential use for smoke control and fire engineering at large. 00:54:24.454 --> 00:54:28.565 So you definitely, definitely should check that episode if you've missed that one. 00:54:29.164 --> 00:54:31.295 And I had an episode with Matt Bonner. 00:54:31.474 --> 00:54:41.985 My friend from Imperial College London, who has also used machine learning algorithms to investigate database of facade fires that we have built together. 00:54:42.434 --> 00:54:50.655 And it was also quite an interesting to see how well the artificial intelligence has carried the task that took such a long time. 00:54:51.105 --> 00:54:54.855 So I'm really, really happy to have this in the podcast portfolio. 00:54:55.405 --> 00:54:58.465 I think these three episodes go together very well. 00:54:58.465 --> 00:55:01.315 And yeah, if you haven't heard them, absolutely. 00:55:01.315 --> 00:55:10.764 After this one, you need to tune in, into Xinyan's episode and Matt's episode, I'm going to drop the links in the show notes And yeah, that's it for today. 00:55:10.855 --> 00:55:13.284 I hope you've enjoyed it as much as I did. 00:55:13.394 --> 00:55:17.864 As usual, next episode, we'll be waiting for you here next Wednesday. 00:55:18.224 --> 00:55:21.164 Looking forward to that and yeah. 00:55:21.195 --> 00:55:21.914 See you around. 00:55:21.945 --> 00:55:22.574 Thank you for listening.