028 - Easy entry into the world of AI in fire with MZ Naser

Fire Science Show

Have you ever been fascinated by the capabilities of AI? Did you wonder how the heck can an algorithm beat humans in repetitive tasks? Or make multi-level correlations that we would never be able to figure out? I was as well. And I felt the urge to learn more about this technology, in a way to not be left out when everyone plays with their new toys... But at the same time, I felt this feeling of overwhelm and confusion about this technology. What exactly is it, where to start... Then the wall of multiple choices to take - am I even trying supervised or unsupervised learning? Is my problem a regression or classification? I won't lie, it's hard already, and I have not even really started yet.

And then comes him. Dressed in white (just kidding). MZ Naser.

MZ is not only a genius who seems to have figured it out in the world of fire, but he is also documenting every step of his path in research papers. More to that, he also wrote a bunch of entry-level papers, and a review paper summarizing the basics and explaining the core concepts. Wow, what a service to the community! Please join me in this discussion in MZ, where he literally walks me through the fascinating world of AI in fire and explains where to start.

At this point of the show notes, I would list you a bunch of papers and relevant resources.
(Update: originally  there  was  just link to MZ’s site, but as this got published you totally need to start with this paper:  https://www.readcube.com/articles/10.1007/s10694-021-01210-1 )

 MZ Naser is so nice, he has a website where all of this is summarized and kept updated! If I started listing the resources, I would do you a disservice... You have to test what he has out there.

https://www.mznaser.com/

And also, please connect with MZ at his Twitter and LinkedIn 

----
The Fire Science Show is produced by the Fire Science Media in collaboration with OFR Consultants. Thank you to the podcast sponsor for their continuous support towards our mission.

2021-11-24 55 min Transcript

Available Results

Generated results are saved to the knowledge database for reuse and search.

No generated results are available for this episode yet.

Extract Knowledge

Pick what you want extracted first. Model, scope, and chapter options appear after a template is selected.

Generated results for public episodes are saved to the knowledge database so they can be reused and searched later.

Transcript

WEBVTT

00:00:00.540 --> 00:00:01.169
Hello, everybody.

00:00:01.169 --> 00:00:03.359
Welcome to Fire Science Show session 28.

00:00:03.839 --> 00:00:07.769
Today we'll be discussing one of my favorite topics in all of the fire science.

00:00:07.769 --> 00:00:09.929
And that is the use of artificial intelligence.

00:00:10.044 --> 00:00:15.833
I consider it one of my favorites because it's something that I would really, really love to learn myself.

00:00:16.089 --> 00:00:21.969
And I'm exploiting this podcast to bring me the best guests that can explain it to me a bit more.

00:00:22.449 --> 00:00:26.640
Honestly, I was quite confused where to start with, , with all of this.

00:00:26.690 --> 00:00:33.320
And after the discussion, as you will hear in this episode, I have a little bit better idea of where to start and how to start.

00:00:33.890 --> 00:00:38.810
And actually I should go on this journey because I'm absolutely convinced it's worth it.

00:00:39.530 --> 00:00:42.710
Today with me I have one of the young leaders of fire safety.

00:00:43.299 --> 00:00:54.820
He's an author of literally countless papers, on AI, in the fire science, I'm actually astonished by the amount and quality of work he's putting through and, publishing.

00:00:55.130 --> 00:00:57.399
I'm really, really admiring him for this.

00:00:57.939 --> 00:01:00.939
Um, he's a professor at Clemson University.

00:01:01.280 --> 00:01:06.003
Yeah, let's just jump into not prolong this because, you want to hear what's after the intro.

00:01:06.003 --> 00:01:34.436
Please welcome MZ Nasser and let's jump into the world of AI and fire science! Hello everybody.

00:01:34.436 --> 00:01:35.846
Welcome to Fire Science Show.

00:01:35.965 --> 00:01:39.956
Today, I'm here with professor MZ Naser from Clemson university.

00:01:40.075 --> 00:01:41.516
Hey Naser great to have you here.

00:01:41.695 --> 00:01:42.266
How about you?

00:01:42.475 --> 00:01:43.016
How's it going?

00:01:43.019 --> 00:01:44.099
I'm fantastic.

00:01:44.099 --> 00:01:45.929
I'm about to learn so much about AI.

00:01:45.929 --> 00:01:49.025
I'm really happy to, or that, Hope so the house.

00:01:49.210 --> 00:01:54.834
I've invited you here because you are a rising star in our industry.

00:01:54.834 --> 00:02:12.281
And you're probably one of very few people who has a good clue about how AI is working and how can we use it, in engineering, actually, there's not so many of them, AI luminaries in our community and, through your papers, which are actually very educative.

00:02:12.311 --> 00:02:17.260
They're not like flashing out you see how advanced algorithms I can use?

00:02:17.260 --> 00:02:25.330
You, you publish a lot of like introductory level AI papers, review papers, and they appreciate that so much.

00:02:25.704 --> 00:02:30.300
What puts you on this pathway to, use computers to enhance your learning?

00:02:30.781 --> 00:02:32.550
This is a very good question, actually.

00:02:32.911 --> 00:02:39.591
So the first time I learned about AI at was, sometime in 2012 or 2013, I was taking a transportation course.

00:02:40.121 --> 00:02:48.021
And in that course, the professor was discussing how we can use AI to organize traffic, traffic, lights, synchronize, all different types of Mm.

00:02:48.111 --> 00:02:48.771
infrastructure.

00:02:49.342 --> 00:03:06.316
And then I kind of did like a very short paper at the time on, fire, on, I get lost because, you know, once you go to your PhD, you focus on experimentations, every simulation of the, uh, you know, those things, then you kind of forget the AI and once I was done, was trying to find the faculty job.

00:03:06.316 --> 00:03:11.445
And as you know, fire tech experiments are very, very expensive and you have to have lab and equipment.

00:03:11.445 --> 00:03:16.346
As you know, you have a massive lab here, in my case, it was very hard to develop this left.

00:03:16.346 --> 00:03:21.536
So I was, I had to do something and I wanted to do something a little bit different than simulation.

00:03:21.536 --> 00:03:33.526
And so I went back to my road to AI and that's when things clicked back again, From 2012 2013, maybe 2019 things has rapidly changed on the AI front.

00:03:33.705 --> 00:03:34.966
Many, many things has changed.

00:03:34.966 --> 00:03:38.145
We have different algorithms, different training systems, learning systems.

00:03:38.716 --> 00:03:41.489
So I had to everything else at that time.

00:03:41.908 --> 00:03:44.729
And hence, some of my papers are just like what you mentioned.

00:03:44.729 --> 00:03:49.929
They are very on the same level because as I was writing them, I was also learning.

00:03:50.468 --> 00:03:51.674
I decided Okay.

00:03:51.758 --> 00:03:53.968
a very smooth way because this is how I learn.

00:03:53.979 --> 00:04:01.598
So perhaps it will also be easier to somebody who was as familiar with AI as me at that time to go about with those papers.

00:04:01.854 --> 00:04:03.504
So, so that's a, that's such a cool path.

00:04:03.504 --> 00:04:08.305
So we basically we're documenting your own ways through the world of AI.

00:04:08.694 --> 00:04:09.985
That's uh, that's so cool, man.

00:04:10.495 --> 00:04:24.081
And that also confirms the theory that you just have to be one step ahead from others, uh, to be an expert and, you know, to, to teach you don't have to know everything to, to provide useful guidance.

00:04:24.081 --> 00:04:26.781
And I really appreciate that you are doing that.

00:04:27.411 --> 00:04:29.545
And I assume on your path to the.

00:04:29.971 --> 00:04:38.742
AI, uh, you've stumbled upon the same confusion everyone is stumbling against for me.

00:04:38.742 --> 00:04:42.045
It's like, What the hell is AI after all?

00:04:42.095 --> 00:04:43.475
Can you even define it?

00:04:43.985 --> 00:04:47.045
And then w what kind of AI should I go?

00:04:47.045 --> 00:04:57.024
Because once you start, like digging, you enter this, uh, loophole with hundreds of algorithms, models approaches, and it's really confusing.

00:04:57.295 --> 00:04:58.314
So, yeah.

00:04:58.504 --> 00:04:59.605
How was it for you?

00:04:59.709 --> 00:05:00.666
Yeah, a hundred percent.

00:05:00.666 --> 00:05:01.505
I didn't know.

00:05:01.956 --> 00:05:03.725
four years ago, I didn't really know.

00:05:03.843 --> 00:05:05.403
or I only knew neuron networks.

00:05:05.583 --> 00:05:07.322
This is what I was Okay.

00:05:07.892 --> 00:05:09.062
it prepped on earlier.

00:05:09.375 --> 00:05:14.016
To me, if this was my math, I could solve anything with neural networks because it was in know Hm.

00:05:14.596 --> 00:05:18.495
one tool that you can button the data points, it would run.

00:05:18.495 --> 00:05:24.136
It should give you some kind of a good performance, if not performance on different problems.

00:05:24.675 --> 00:05:29.115
as you mentioned, nowadays, we have all these different types of learning or these of algorithms.

00:05:29.115 --> 00:05:32.326
And the easiest thing in my case was I need to learn.

00:05:32.326 --> 00:05:32.716
I need to learn.

00:05:33.451 --> 00:05:37.560
So I had to go back and see, okay, what are the basics for supervised learning?

00:05:37.560 --> 00:05:39.060
What, what is classification?

00:05:39.060 --> 00:05:45.911
What is regression and , so once you go back to computer science and see those definitions, then you would see, okay, know what?

00:05:45.911 --> 00:05:49.451
Most of the problems we do in fire engineering are really regression.

00:05:49.661 --> 00:05:50.990
You know, we have a phenomena.

00:05:51.446 --> 00:05:54.985
And the outcome of this phenomenon is a number in a fire resistance.

00:05:55.196 --> 00:05:58.586
It could be like, you know, heating, great burning, great, some kind of a number.

00:05:59.185 --> 00:06:14.586
you're out to some kind of a number, then this is a very good chance that you are going to be dealing with a supervised learning problem with a sub a component that's going to be a regression if your output is going to be something like a category, for instance, this column fails or doesn't fail, slap collapse, it doesn't collapse.

00:06:14.946 --> 00:06:16.935
You have charring and you don't have charring.

00:06:17.646 --> 00:06:25.103
instance, this fire is heat, uh, in a ventilated control, you are trying to put the phenomenon into one group, this is classification.

00:06:25.523 --> 00:06:36.862
So once you know the problem, you have to define the algorithm, then you'd say, okay, well now my problem is for instance, regression, what kind of algorithms are there out there now that can solve a regression problem?

00:06:37.360 --> 00:06:42.029
So from there you will go, you'll find hundreds of algorithm that can do the same.

00:06:42.819 --> 00:06:46.509
So the question becomes what, which one of these algorithms I'm going to, I'm going to use.

00:06:46.540 --> 00:06:54.860
And the answer is to be honest, you could potentially use any single algorithm of these and if you have a good database, you will come up with a good answer.

00:06:55.040 --> 00:07:06.870
You would come up with a good prediction to the problem becomes, as you might have guessed is why would I go with algorithm A, instead of algorithm, B or C or D what are the motivation behind these algorithms?

00:07:07.505 --> 00:07:13.386
the answer to this is interesting because this is exactly like saying, shall I use ANSYS or Abacus to solve a problem?

00:07:14.055 --> 00:07:16.565
It's basically which algorithm you're familiar with.

00:07:16.625 --> 00:07:17.406
It's the company.

00:07:17.507 --> 00:07:19.017
To your own experiences.

00:07:19.017 --> 00:07:20.637
In my case, I've always used ANSYS.

00:07:20.668 --> 00:07:23.067
I use Abacus very, very slightly.

00:07:23.067 --> 00:07:26.697
So if you go back to my papers, they're all ANSYS, thing with my algorithms.

00:07:26.697 --> 00:07:29.997
You'll see that the earlier work was heavily towards neural network.

00:07:30.867 --> 00:07:40.257
recently, I've learned more about different algorithms, the modern ones, because now as you know, more than algorithms are almost superior when it comes to trying to for prediction power.

00:07:41.098 --> 00:07:54.160
to be honest, if you have a nice database, if you run, let's say 10 algorithms out of the 10 algorithms, most likely nine of them will give you R square or of 95%, 90%, 85%.

00:07:54.819 --> 00:08:01.343
it's the science is really not in running the algorithm, the sciences what did you learn from this algorithm?

00:08:01.822 --> 00:08:03.382
let's say that you use this algorithm.

00:08:03.382 --> 00:08:05.589
You have a good performance, but how does this.

00:08:05.697 --> 00:08:18.446
Actually advances our science or our knowledge, So to break the first wall for anyone jumping into, , I, from your papers I've learned or supervised, unsupervised and semi-supervised methods.

00:08:18.466 --> 00:08:24.543
And, this seemed like the very first critical choice, uh, one would make, uh, when they enter.

00:08:24.543 --> 00:08:37.009
So could you like try and briefly showcase the differences and,  and give these examples, like So supervised learning term supervise means, you know, the inputs you know, the output.

00:08:37.129 --> 00:08:38.159
So everything is being.

00:08:38.784 --> 00:08:43.559
Um, So for instance, let's say that we are trying to figure out if this column is going to fail under fire.

00:08:43.919 --> 00:08:47.009
We know the column, geometry, we know its material properties.

00:08:47.009 --> 00:08:51.600
We know if it's going to be boundary conditions fixed, wing, all of these things, we've done a test.

00:08:51.690 --> 00:08:57.720
So we also know it's fire resistance, or we know when it's going to fail, you know, everything, you know, the inputs, you know the output.

00:08:57.730 --> 00:08:58.620
So this is supervised.

00:08:59.190 --> 00:09:07.472
Let's say now, know, all the inputs, let's say you have a group of, columns, you know, all their inputs, but you don't know when they fail.

00:09:08.269 --> 00:09:15.645
So you would use unsupervised learning and this way the algorithm should cluster or combine the columns that are similar to each other into groups.

00:09:15.645 --> 00:09:22.696
And then the algorithm would say, well, these five columns are group one, these four columns or group two, these four columns are group three.

00:09:22.905 --> 00:09:24.405
You don't know the output.

00:09:24.870 --> 00:09:34.490
You don't know why these are in groups, but if you go back and study the fire test results, you're likely to see that the columns and the group one, maybe they failed within an hour.

00:09:34.809 --> 00:09:35.868
Group two maybe they Okay.

00:09:35.942 --> 00:09:36.753
within two hours.

00:09:37.352 --> 00:09:41.072
unsupervised learning is when you know the inputs, but you don't know the output.

00:09:41.082 --> 00:09:42.543
You don't know what the phenomena is.

00:09:42.602 --> 00:09:43.982
You're just trying to group things together.

00:09:45.062 --> 00:09:47.462
Semi-supervised learning is going to be somewhere in between.

00:09:47.462 --> 00:09:53.822
Sometimes semi-supervised would be something that say that we have, images of columns failing.

00:09:54.452 --> 00:10:00.243
instead of us going by image and saying this column, fails this column doesn't fail.

00:10:00.572 --> 00:10:10.452
We could potentially only label 50 images and the algorithm should be able to label the additional 50 images that we did in labor.

00:10:10.452 --> 00:10:13.302
So this way it has a little bit of knowledge on the inputs outputs.

00:10:13.753 --> 00:10:15.373
doesn't have it for all the database.

00:10:15.972 --> 00:10:19.212
So it's somewhere in between supervised and unsupervised learning.

00:10:19.772 --> 00:10:37.635
So, if you had a, let's say a supervised algorithm with the database on existing columns, and then you come up with a completely new column, the supervise would tell you when it will fail, based on its knowledge, the unsupervised would, it would tell you to which group of columns this one looks more familiar to.

00:10:37.635 --> 00:10:42.706
And the semi-supervised could just continue the task you were doing with the previous columns.

00:10:43.035 --> 00:10:47.655
Was it painting it pink or measuring their moment of inertia or something?

00:10:47.655 --> 00:10:48.105
Okay.

00:10:48.775 --> 00:10:49.676
This seems useful.

00:10:50.123 --> 00:10:50.633
Nice.

00:10:50.663 --> 00:10:53.349
Uh, reminded me when you said, your technician has an AI.

00:10:53.589 --> 00:10:58.678
This is what his algorithm would do, algorithm would recognize noise, or maybe like my mic the hoodie.

00:10:59.158 --> 00:11:12.089
it would label that as this is like noise or this is like not voice because it has seen before through training that this sign of scratching is not really a voice, so you have to take it away.

00:11:12.535 --> 00:11:22.885
let's just jump quickly from enthusiasm to the dangerous region, because you've mentioned it seen, but if has not seen something, it's very unlikely.

00:11:22.885 --> 00:11:24.926
It's going to predict the behavior, right?

00:11:25.166 --> 00:11:31.916
Like if you, if you show it a thousand fires with flashover, it will not know that backdraft my may happen.

00:11:31.916 --> 00:11:32.155
Right.

00:11:32.765 --> 00:11:33.696
And this is the problem.

00:11:33.905 --> 00:11:35.135
This is exactly the problem.

00:11:35.166 --> 00:11:53.105
The problem is when you develop an algorithm and have a good database and you have good performance, the researcher or need to remember that this performance is only valid for your database to go beyond the database is going to be very, very tricky because when you have a database, you're immediately constraint in your algorithm.

00:11:53.115 --> 00:11:55.841
So you have a space of, oh, you have a okay.

00:11:56.015 --> 00:11:56.975
a space of inputs.

00:11:57.446 --> 00:12:01.765
You can possibly collect everything and you can collect some features of the space.

00:12:02.186 --> 00:12:07.166
And then for the algorithm on what it sees is features as the whole space.

00:12:07.346 --> 00:12:10.602
So if you have an additional feature outside of this space.

00:12:11.023 --> 00:12:13.778
It's going to be very hard to give you a correct production.

00:12:13.778 --> 00:12:18.399
Maybe it could sometimes if, if the algorithm or maybe if the problem is simple enough, it could.

00:12:18.879 --> 00:12:23.739
But other than that is going to be very tricky that's very similar to experience, actually.

00:12:24.009 --> 00:12:24.249
Yeah.

00:12:24.349 --> 00:12:29.139
If you experienced a lot of things, you're more likely to predict things.

00:12:29.239 --> 00:12:33.105
That's uh, that was something we share with the machines, I guess.

00:12:33.138 --> 00:12:33.317
Yeah.

00:12:33.317 --> 00:12:43.780
W one thing, yeah, experience, value this a lot when we have experienced, usually like at least us, we have a knowledge of what could happen.

00:12:43.811 --> 00:12:51.051
Like we could see beyond that we have algorithms can't and that's, that's going to be the problem that we're going to be dealing with.

00:12:51.270 --> 00:12:54.380
We can go beyond what we can see beyond the data.

00:12:54.620 --> 00:12:58.490
However, very, very good to see between the lines and we're not.

00:12:59.030 --> 00:13:06.380
So this is how we can compliment both of us, because if you have a complex database for us, it's going to be very hard to visualize for them.

00:13:06.380 --> 00:13:06.890
It's easy.

00:13:06.890 --> 00:13:15.073
If it can see things, and this is why they predict things with high accuracy, but that doesn't mean that this prediction is actually something that's physically correct.

00:13:15.645 --> 00:13:44.566
I had this episode on AI and fire already with Xinyan Huang from Hong Kong Polytechnic University, and Xinyan is doing a lot of crazy things with, , with smoke control fire detection in tunnels and you also mentioned that this, human machine, uh, combination is the most powerful  and in a way I had a feeling he would like this AI be a way you could transfer the collective experience of whole industry.

00:13:44.936 --> 00:13:51.316
And that for me, that was such a powerful and beautiful idea that, , so much knowledge is lost between us.

00:13:51.316 --> 00:14:01.546
And if we could have this collective mind helping each other, it would be fun, but it's also seems very difficult from the technical point of view to achieve that.

00:14:01.546 --> 00:14:01.875
Right.

00:14:02.296 --> 00:14:14.298
Because, like, to what extent the, um, structure of the database is also important, like to what extent you can drop scattered data into an algorithm and expect correct results.

00:14:14.817 --> 00:14:14.899
Okay.

00:14:15.533 --> 00:14:17.634
first of all, we're not computer scientists.

00:14:17.634 --> 00:14:20.139
were appliers.Computer Yeah.

00:14:20.224 --> 00:14:24.173
the algorithms, they validate them over multiple databases.

00:14:24.173 --> 00:14:24.953
We just take them.

00:14:24.953 --> 00:14:28.553
And then we do our own little experiment and we have good performance and we think it works.

00:14:29.283 --> 00:14:46.293
second part of the issue as the machine learning we're using now, or the algorithms that we're using now, they're highly data driven or correlations driven which negates the purpose of science and not everything correlates that there is a cause and effect.

00:14:46.803 --> 00:14:48.063
This is why, Okay.

00:14:48.244 --> 00:15:02.494
myself, I'm trying to move away from all this data driven nonsense and go towards like modern algorithms that at least can give you a cause and effect because if you know the cause and effect, but regardless of how much data you have, always get the right answer.

00:15:02.644 --> 00:15:04.384
The goal is to know why this.

00:15:05.283 --> 00:15:12.874
That goes, the issue is not to know seeing this, or I have seen this in 10 experiments and this would happen in the 11th extrovert experiment.

00:15:12.884 --> 00:15:13.724
There is no guaranteed.

00:15:14.203 --> 00:15:15.403
Observations help.

00:15:16.004 --> 00:15:19.844
However, to come up with knowledge, you need to know why, and it's cause and effect.

00:15:20.323 --> 00:15:32.203
And if if you teach an algorithm cause and effect, then you have to completely negate or move away from the type of learning that we have now in commercial machine learning, commercial machine is purely data driven.

00:15:32.258 --> 00:15:35.768
Now I'm not saying that correlation or data driven doesn't have a purpose.

00:15:35.768 --> 00:15:39.187
I do have a, but it does have a purpose and it would work for different problems.

00:15:39.937 --> 00:15:55.687
for our own, if you want to advance knowledge, as opposed to apply knowledge, is going to be fine for correlation or data driven, because you're looking for a solution you want, see this every day, you want a surrogate model that tells you if you see this, this is likely to happen.

00:15:55.687 --> 00:15:55.988
You know what.

00:15:56.783 --> 00:16:01.043
But if you want to know why things happen, can't rely on AI.

00:16:01.043 --> 00:16:08.753
We have to combine AI into our experiments and we have a completely different kind of teaching methods for AI to figure out cause and effect.

00:16:09.182 --> 00:16:21.452
In one of your papers or in one of your talks you've used in definition of AI as a computational technique that exploits hidden patterns between seemingly unrelated parameters to draw solutions to a given phenomenon.

00:16:21.812 --> 00:16:28.393
But often when you see these data driven AI it just seems like really complex statistics.

00:16:28.393 --> 00:16:35.980
You know, it's like something you could not, plot and having the R square on a single plot is drawn from multiple dimensions, let's say.

00:16:36.519 --> 00:16:39.268
And, This statistical one, it seems attractive.

00:16:39.268 --> 00:16:49.168
It's interesting, possibly useful and probably very useful, but it's this exploitation of hidden patterns between seemingly unrelated parameters.

00:16:49.528 --> 00:17:00.668
This seems like something that could tell us why the facades are burning or why spalling occures, or I don't know why in some conditions, firefighters may die  in, in the room.

00:17:00.967 --> 00:17:09.127
So to, uh, but to achieve these hidden pattern uh, recognition, you need knowledge beyond data, right?

00:17:09.127 --> 00:17:11.238
You need to have observations.

00:17:11.238 --> 00:17:13.548
That's what you meant by coupling the experiment and AI AI.

00:17:13.653 --> 00:17:17.403
need to have, you need to have a methodology of  saying this.

00:17:18.182 --> 00:17:19.083
have experiments.

00:17:19.083 --> 00:17:23.643
I've seen this, but this experiment is going to be limited by whatever equipment I have sensors.

00:17:23.643 --> 00:17:24.063
They have.

00:17:24.063 --> 00:17:25.885
So sometimes I'm picking up data.

00:17:25.915 --> 00:17:26.786
I think it's noise.

00:17:26.786 --> 00:17:27.776
Maybe it's not noise.

00:17:28.135 --> 00:17:31.286
So you have to do a multiple levels of experimentation.

00:17:32.155 --> 00:17:38.415
Use that data and you have a teacher algorithm at different level what each one means.

00:17:38.415 --> 00:17:45.718
And then the algorithm should be able to put an overall picture of hidden weight pathways between how these factors react.

00:17:46.167 --> 00:17:57.317
For instance, many research papers now on AI and and not just in the fire in really any, any field in engineering, the first, the second user, the second section of a paper but it would be like description of database.

00:17:57.377 --> 00:17:59.278
And then they would list database.

00:17:59.298 --> 00:18:04.278
And then they would say all these that we have min max average median for this, for our database.

00:18:04.817 --> 00:18:15.678
And the second thing that always worked is like a correlation matrix and then they would say, this is the correlation between, the correlation matrix is only going to be linear because you're using a linear correlation.

00:18:15.887 --> 00:18:23.778
There is no guarantee the relationship between the factors themselves or the features themselves or the features and output is linear.

00:18:24.438 --> 00:18:28.644
So having that metrics or that table doesn't really tell me anything about cause and effect.

00:18:28.644 --> 00:18:32.855
It tells me that I could use, any statistical model to come up with an equation.

00:18:33.924 --> 00:18:35.914
the thing about machine learning and statistics is the.

00:18:36.769 --> 00:18:47.059
The statistics are too, as a model you have to be confined with the assumptions of that model because model that has a certain kind of assumptions, applicable for a set of data.

00:18:47.539 --> 00:18:54.380
And hence the statistic can't be applied to many, many of our problems because we have highly on the know problems in machine learning.

00:18:54.380 --> 00:18:59.039
These are non-parametric so they don't have for data.

00:18:59.220 --> 00:19:00.960
We don't have assumption for distribution.

00:19:01.140 --> 00:19:14.055
They can fit very complex functions within our database, most likely they outperform statistics in our case, our problems were linear, you won't find anybody using machine learning because why would you use machine learning Yeah.

00:19:14.140 --> 00:19:15.579
a comfort, such a simple problem.

00:19:16.111 --> 00:19:23.250
I guess in 20 or 30 years, when this becomes like mainstream, you will see people using machine learning for linear problems.

00:19:23.590 --> 00:19:29.240
Just like today, we use a CFD for extremely simple cases instead of zone models.

00:19:29.280 --> 00:19:30.990
I predict is going to happen.

00:19:31.411 --> 00:19:51.799
And, you didn't say that, but I assume that it's also highly related to the amount of data you have to be able to get this quality or this multilevel correlations there, and, like how much that's w when does one know they have enough data for the problem.

00:19:52.309 --> 00:19:58.869
And and another question that's something I would be very interested in when I am planning my experiment.

00:19:59.540 --> 00:20:04.313
How should I prepare myself to create sufficient amount of data, you know?

00:20:04.462 --> 00:20:14.032
So I want to have a grant and I need to know if I will need a thousand experiments, a hundred experiments or 10 experiments of this type and 50 additional of another type.

00:20:14.717 --> 00:20:16.276
Yeah, it's, it's a very good question.

00:20:16.306 --> 00:20:20.476
And it's one, I know there is a lot of research in computer science to figure out this answer.

00:20:20.955 --> 00:20:29.415
know there are a few papers, like I think the 5, 5, 6 years ago, the minimum number of observations somebody would need is maybe I think 10 or 12.

00:20:30.016 --> 00:20:31.605
Now the number is, is 25.

00:20:31.605 --> 00:20:38.836
So if you have a 25 observation, most of the time, you should be able to have some kind of a model that performs in a nice way.

00:20:39.277 --> 00:20:48.506
I know, however, for some type of learning classification problems, think the minimum number that you could be confident about your results is about a hundred observation.

00:20:48.977 --> 00:20:52.369
So these are the numbers that we play with, somewhere between 25 and a hundred.

00:20:52.369 --> 00:21:07.829
Now, the problem becomes is the database becomes wider, you have many, many features, then you will have to have many observation, because if you think about it, a database is a matrix have rows and column, the wider, the matrix so you have to have very, very deep.

00:21:08.349 --> 00:21:11.140
Figured out some kind of correlation between the different parameters.

00:21:11.559 --> 00:21:15.039
So the wider it is the more data that would need.

00:21:15.609 --> 00:21:18.039
And I don't have the, I don't know the answer.

00:21:18.039 --> 00:21:19.930
I don't know if you have 50 data points.

00:21:19.930 --> 00:21:22.269
Is this going to be enough for no, I really don't know.

00:21:22.269 --> 00:21:38.500
I don't think we'll have an answer anytime soon, if we are within, 25 to a hundred points with maybe four to seven features, we should be, or at least the algorithms we have now would be able to give you something that's, maybe with some confidence.

00:21:38.920 --> 00:21:48.279
And to compliment what I just mentioned, , nowadays the algorithms that we do use, they could be augmented with, uh, different tools.

00:21:48.279 --> 00:21:50.940
For instance, you can add confidence intervals to your model.

00:21:51.420 --> 00:21:56.819
So this way, even if you have a small database, algorithm should to the, okay, this is my production.

00:21:56.819 --> 00:22:00.099
I'm predicting this column to fail in 16 minutes.

00:22:01.049 --> 00:22:06.075
this is going to be within a confidence of 90% or 70% of 60%.

00:22:06.404 --> 00:22:12.290
So even if you have a short database, you'd have a prediction, but on the other hand, you have some kind of confidence in your model.

00:22:12.320 --> 00:22:22.040
It's not like the old days, two, three years ago when we couldn't apply confidence intervals were more, and it just, you know, this is the number that you have We could potentially add confidence.

00:22:22.040 --> 00:22:33.740
And then this would give us some level of trust because even if we have a short database or a very, very wide database, that prediction is going to be combined with some kind of confidence that would say, okay, my prediction is this much with 70%.

00:22:34.398 --> 00:22:55.127
In the past, I was, maybe I was not working much with them, but I, I got familiar a bit with some design of experiment methods that allow you to, , figure out the number of experiments to identify the, for example, the influence  of variables, on the outcome, Benheken  design, , there was aLatin Hypercube, something, Monte-Carlo of course, uh, many, many methods like that.

00:22:55.518 --> 00:23:02.478
And, uh, I was very interested in them because in my PhD, I did, , a very ugly thing.

00:23:02.988 --> 00:23:09.238
I've taken like a hundred geometries,  a few fires and some combinations of ventilation.

00:23:09.238 --> 00:23:10.647
And I just brute force to them.

00:23:10.647 --> 00:23:16.498
And it, it gave me a beautiful array of results to work with and complete my PhD.

00:23:16.498 --> 00:23:17.968
And I was very happy with it.

00:23:17.968 --> 00:23:19.917
I'm still, I still am happy with that.

00:23:20.307 --> 00:23:29.394
I just feel like,  caveman bashing, a wooden stick against a wall to get an answer where I could have done this way more elegant in a way.

00:23:29.744 --> 00:23:36.404
So, so is there also let's say preparation best practices to drive an experiment.

00:23:36.404 --> 00:23:37.724
So it's useful for this.

00:23:38.387 --> 00:23:59.468
So there are a few things, for instance, we have some algorithms that what all, what they do is they look at the distribution of your observations or the experiments that you have so far, then they are able to zoom or to pinpoint some regions within that distribution that say, okay, this region or for this distribution, we need to have more experimental data points.

00:23:59.468 --> 00:24:01.248
So this way, let's say that Okay.

00:24:01.448 --> 00:24:02.347
to three experiments.

00:24:02.768 --> 00:24:03.518
have to keep in mind.

00:24:03.518 --> 00:24:06.938
Maybe I need to allocate maybe two more experiments with this.

00:24:06.968 --> 00:24:11.567
That would give me To cover a specific, uh, variable for example.

00:24:11.567 --> 00:24:11.928
Okay.

00:24:12.008 --> 00:24:19.738
Because if you think a model and how would validate itself, basically runs some kind of, performance metrics along different regions.

00:24:20.458 --> 00:24:25.008
we normally do is we collect all the data points or the predictions, and we run R square.

00:24:25.587 --> 00:24:26.907
Now R square doesn't tell you.

00:24:27.400 --> 00:24:29.462
the, the performance for specific points.

00:24:29.462 --> 00:24:32.313
It tells you the performance for the whole database or for the whole predictions.

00:24:32.823 --> 00:24:40.502
But if you plot X and Y would see that your curve at some regions are much, much larger than the other regions.

00:24:40.712 --> 00:24:46.202
And for those regions, you might want to end up with experimental points because the distance is two things.

00:24:46.563 --> 00:24:50.883
This one's you want, the algorithm was not able to capture that phenomena at that region.

00:24:50.913 --> 00:24:54.752
And two, it tells you that there is something there that we haven't seen before.

00:24:55.173 --> 00:25:01.653
So maybe if you do an experiment, maybe you'd be able to confirm it or deny it and figure out something new.

00:25:01.726 --> 00:25:08.942
So this will guide you towards the potential outlier that could actually unravel a new physics or something completely unexpected.

00:25:08.942 --> 00:25:09.163
That's.

00:25:09.163 --> 00:25:09.752
That's cool.

00:25:10.272 --> 00:25:13.982
Okay, let's move a bit more into engineering.

00:25:14.712 --> 00:25:33.160
I I've taken a look on your papers and you have used AI to identify fire vulnurable bridges, designed columns, change measurements, for fiber reinforcement, polymers, strengthens reinforced columns to determine spalling it identify failure of beams.

00:25:33.490 --> 00:25:38.289
Is there any field of engineering you have identified that it will not work at all?

00:25:38.559 --> 00:25:39.009
Maybe?

00:25:39.250 --> 00:25:42.099
this is a no, it's a, it's a very good question.

00:25:42.480 --> 00:25:49.720
And then this is what scares me most, to be honest, I, I'm a little bit lucky because I get to play with AI a little bit earlier Yeah.

00:25:49.859 --> 00:25:50.789
to see how it works.

00:25:51.240 --> 00:26:08.303
And so far it's working really well, which tells me it's either the problems we have are enough for AI, because if you really think about it, scientists, when they develop an algorithm it works for insurance, for medicine, for space, it's not just, know, bending.

00:26:08.333 --> 00:26:10.313
Buckling, flammability, collapses.

00:26:10.792 --> 00:26:12.323
have much, much more complex problems.

00:26:12.833 --> 00:26:19.086
if it works well for complex problems, like finding a new star galaxies, that's a very complex problem.

00:26:19.566 --> 00:26:29.893
Maybe our problems not that after all, for AI to solve, maybe the, maybe it's complex for empirical methods that we apply or for finite elements methods that we use.

00:26:30.002 --> 00:26:31.063
They're not really that hard.

00:26:31.063 --> 00:26:34.093
They're just computationally expensive for FEA or like FDS.

00:26:34.153 --> 00:26:35.682
You have to run them for a long time.

00:26:35.682 --> 00:26:41.442
You have to mention it had elements, but collectively, maybe they're not the other issue I'm thinking.

00:26:41.442 --> 00:26:48.455
I mean, I think about is maybe because algorithm is really a black box and we don't know why, how it behaves.

00:26:48.873 --> 00:26:49.682
We only see the.

00:26:50.522 --> 00:26:56.252
What if the output is correct, but the map that links the inputs to the output is not correct.

00:26:56.762 --> 00:27:04.442
Maybe the output is correct for this database, but once you, once you go for outside of the range of database, it's going to be very, very hard.

00:27:04.472 --> 00:27:09.182
Or maybe you start to get some errors that the counterpart is the following.

00:27:09.833 --> 00:27:16.613
The counterpart is very interesting because let's say in structural fire engineering, like we have certain sizes for columns and beams.

00:27:17.103 --> 00:27:18.123
don't go beyond that.

00:27:18.807 --> 00:27:19.127
Hm.

00:27:19.173 --> 00:27:24.663
databases are usually good because, you won't find a very, very, very thin column or a very, very short column.

00:27:24.663 --> 00:27:25.413
We don't use that.

00:27:26.042 --> 00:27:31.836
going outside the norm will give you error measurements, but at the same time, we never used in practice.

00:27:32.496 --> 00:27:40.654
there is like a and cons for, for every issue In the same way, uh, you will have only combustion within limits of flammability.

00:27:40.924 --> 00:27:45.454
You would have certain sizes of fire only in certain, uh, ventilation factors.

00:27:45.785 --> 00:27:51.724
So there are like boundaries to the fires that we know empirically and we could work with that.

00:27:51.922 --> 00:28:03.288
With all these innovation that you show in machine learning, I'm really wondering how hard is it in a field so let's say a field with concrete, like ours.

00:28:03.528 --> 00:28:04.877
Like, uh, construction.

00:28:05.057 --> 00:28:07.877
It's not a place of raging innovation.

00:28:07.968 --> 00:28:15.471
It's, we're using hundred years old standards to quantify a fire resistance and it's not unlikely to change very soon.

00:28:15.891 --> 00:28:18.681
The problem is not with innovating something.

00:28:18.891 --> 00:28:29.171
The problem is innovating something and not breaking, everything else, and we're very slow to adopt new technologies, new methods, new.

00:28:29.330 --> 00:28:34.010
How is it going for you as a pioneer of this technology in construction?

00:28:34.317 --> 00:28:36.627
You must have a funny reviews for, for your papers.

00:28:37.835 --> 00:28:42.275
Um, papers, I would get very unique reviews.

00:28:42.275 --> 00:28:42.724
Yes.

00:28:43.085 --> 00:28:58.085
I would say now it's much better now and it much, much better, I would say the following a hundred percent with a slow to adapt, I would agree with this, like two, two years ago, when I was trying to push for something in a conference or with a funding agency, the answer is not complete.

00:28:58.144 --> 00:29:03.244
So I had to reconsider my whole path because I couldn't get anything from anybody.

00:29:03.875 --> 00:29:06.434
And nowadays I think the industry is interest.

00:29:07.123 --> 00:29:10.685
Like we, we had, with the American concrete Institute, we had a few talks.

00:29:10.685 --> 00:29:13.685
We have, we published a book with them on AI, completely with concrete.

00:29:13.685 --> 00:29:15.125
They were very, very open to it.

00:29:15.201 --> 00:29:15.671
okay.

00:29:15.806 --> 00:29:24.486
the steel industry is looking to something in a close by mass and we actually, one of my grants is from masonry they want to use AI to design masonry structure for fire.

00:29:25.415 --> 00:29:30.185
it's a little bit bad, but that the thing that's good for our case right now is the following.

00:29:30.546 --> 00:29:33.349
We have a lot of startups many of these are startups.

00:29:33.380 --> 00:29:38.869
They're trying to automate many of the routine applications or that we use.

00:29:39.440 --> 00:29:46.009
then these are a way to automate our routine step is to use machine learning because you know, you're not really going above and beyond in something.

00:29:46.339 --> 00:29:48.170
You're you have a procedure already.

00:29:48.170 --> 00:29:54.920
You're just trying to make it much, much faster, much more accurate, less error, all of that would accumulate to less time, more money.

00:29:55.700 --> 00:30:01.460
So there is a push from the industry now, and I think it will grow within the next two or three years because there's a lot of startups.

00:30:01.970 --> 00:30:02.750
these are startups.

00:30:02.779 --> 00:30:04.099
If you look at the investors with.

00:30:05.009 --> 00:30:06.230
They're not really engineers.

00:30:06.289 --> 00:30:10.339
They're mainly into developing softwares and apps and computer science backgrounds.

00:30:10.940 --> 00:30:13.940
however, they don't have the domain knowledge that we do.

00:30:13.940 --> 00:30:17.930
And this is why they hire civil engineers or structural engineers or fire engineers.

00:30:18.680 --> 00:30:22.202
they learn the problem, it's going to be very, very easy for them to develop solutions.

00:30:22.232 --> 00:30:24.752
The issue with me is the following.

00:30:24.752 --> 00:30:32.836
The issue is we need solutions that come from somebody who has been practicing and been educated in our field solve our problems.

00:30:33.316 --> 00:30:40.185
We can just give the domain knowledge to somebody who's doesn't have our background because they're looking at the surface.

00:30:40.185 --> 00:30:44.553
We need people to do a fire engineering or structural engineering from the beginning, from the under.

00:30:45.288 --> 00:30:51.887
Then you would come up with solution that would work better, would work best for our case and will advance our knowledge as well.

00:30:51.917 --> 00:30:56.657
Because you know, I'm not really looking for a software that tells me is the amount of fire you're going to get.

00:30:56.657 --> 00:31:00.567
Or like, this is the heat intensity you're going to get, because anybody can do a software like this.

00:31:00.807 --> 00:31:04.857
I want to know why, if I know why redesign, I can change things.

00:31:04.857 --> 00:31:09.087
I can come up with unique designs, innovations that we don't have right now.

00:31:09.478 --> 00:31:12.928
And the computer scientists can give you that you have to be an engineer of that.

00:31:13.444 --> 00:31:20.065
I'll challenge that because,, for a paradigm shift to occure in a field, it must be done by someone from outside, outside of that field.

00:31:20.144 --> 00:31:28.194
If you're a graduated fire safety engineering, it's very unlikely that you will change the fire safety engineering completely because of the way how you would have been thought.

00:31:28.194 --> 00:31:36.414
And there is certain experience factor in your computer in your head that, uh, will prevent you from touching stuff.

00:31:36.444 --> 00:31:47.934
But our, you escaping the field and jumping into computer science and coming back is actually quite a nice path to, to carve such a path for for new, uh, if also you stern black books many times.

00:31:47.934 --> 00:31:50.318
And, it seems like that, I don't understand it.

00:31:50.318 --> 00:31:54.459
You know, I see, I know I can put some stuff in, it will give me stuff out.

00:31:54.701 --> 00:31:55.781
Even for CFD.

00:31:55.781 --> 00:32:01.011
I'm in CFDs very, very hard, but I can more or less understand CFD.

00:32:01.031 --> 00:32:01.602
I don't claim.

00:32:01.602 --> 00:32:06.912
I understand it completely, but I, more or less know what the equations do, what the schemes are.

00:32:06.922 --> 00:32:09.912
What's a turbulence model, what's boundary layer.

00:32:09.942 --> 00:32:15.672
I know these things and I can track back my simulation, identifying each of these steps and going in.

00:32:16.001 --> 00:32:19.842
And then I see a pattern of neural network and it looks like a Christmas tree to me.

00:32:20.082 --> 00:32:23.412
It doesn't reassemble a anything equation or something.

00:32:23.412 --> 00:32:33.082
And, in your paper, in, um, Automation in Construction, engineers guide to AI, you you've championed this explainable AI as a necessity.

00:32:33.692 --> 00:32:41.461
So tell me what would be this explainable AI and why it would be something that would make me use AI and while I'm not using it today.

00:32:41.852 --> 00:32:42.031
Yeah.

00:32:42.632 --> 00:32:45.122
So what I did is it's, it's a very simple exercise.

00:32:45.122 --> 00:32:54.842
So you take an equation from a code and you apply in a database and you see that the equation from the code that we have to use as engineers does not perform as well as an algorithm.

00:32:55.321 --> 00:32:58.082
So this fact by it sort of should make you pause.

00:32:58.082 --> 00:33:01.471
How can you not trust a code over an algorithm?

00:33:02.281 --> 00:33:13.204
Then the second question would be if the algorithm can predict better than the code, why do I have to use the code when they have a better method that can predict better than the code?

00:33:13.805 --> 00:33:15.541
The second question is the following.

00:33:16.172 --> 00:33:18.541
Why does the algorithm do that?

00:33:19.051 --> 00:33:19.652
the code cannot.

00:33:20.691 --> 00:33:23.277
Now to know why an algorithm does a certain thing.

00:33:23.277 --> 00:33:27.567
We have to break that, and see how it, how it does the way it does.

00:33:28.017 --> 00:33:34.136
And right now we don't know we can do that because one would not commit a scientist even computer scientists.

00:33:34.166 --> 00:33:44.517
Can't really track how the algorithms work, because everything for them is goal is to get as good of a prediction as an experiment, or as an observation, don't care how to get there.

00:33:44.517 --> 00:33:47.636
In our case, we do care because we have to justify our decisions.

00:33:47.936 --> 00:33:52.977
How can they justify using a column with two hour fire rating in the building they don't know why?

00:33:52.977 --> 00:34:03.297
If the algorithm says, yes, I, if I know why need to figure out what, so this is when, do you have to use If you have AI, you have also above it, explainable AI, explainable AI.

00:34:03.866 --> 00:34:05.096
The way it does is the following.

00:34:05.186 --> 00:34:09.896
Each algorithm should be able to tell you exactly how it came up with its own production.

00:34:10.527 --> 00:34:11.996
It has to break it down for you.

00:34:11.996 --> 00:34:15.699
So you can understand because numerically It's correct.

00:34:15.699 --> 00:34:19.960
However, physically, or from an engineering maybe it's not correct.

00:34:20.679 --> 00:34:27.494
the algorithm of say the relationship between, you know, material and geometry is, is linear, but we know from our experiment, it's not linear.

00:34:27.943 --> 00:34:29.293
how can I trust it to production?

00:34:29.344 --> 00:34:31.384
If it negates what physics tells me to do.

00:34:32.074 --> 00:34:34.534
Now, the problem with explainable AI is the following.

00:34:34.976 --> 00:34:39.056
It will only explain its results based on the database that you have.

00:34:39.117 --> 00:34:53.067
So if you don't have a good database or as many features as the physics would allow you to do, even if you have explainable AI, it won't be as good as the one we have in physics, because it's, won't be able to capture all the interactions that one we see in physics.

00:34:53.367 --> 00:34:56.246
So it's not really about using explainable AI or AI.

00:34:56.246 --> 00:34:59.786
It's about using a system that can tell you this because.

00:35:00.432 --> 00:35:07.956
When we use explainable AI, we basically have a very small code within our algorithm that can track prediction back to its origin.

00:35:08.527 --> 00:35:14.893
did the algorithm link, parameter one with parameter two with parameter three with parameter, for to come up with a prediction that it did.

00:35:15.117 --> 00:35:15.956
Black box.

00:35:15.987 --> 00:35:41.693
Doesn't tell you that So for example, if I employed, AI to predict, , smoke movement in a, let's say buoyant plume, it could actually, in the meantime, tell me that it works when you assume the gravity is less on the Mars and then it works while in fact, it's just a matter of entrainment coefficient that is elsewhere with which could accidentally be the same number as the ratio of gravity here in the Mars.

00:35:42.023 --> 00:35:44.934
But then the algorithm will never know what happened.

00:35:44.934 --> 00:35:46.824
It just used this and it worked.

00:35:46.824 --> 00:35:49.360
And for them, it's, perfect for engineering.

00:35:49.360 --> 00:35:58.260
You need to understand, and this, uh,  so breaking it into steps and seeing the more or less what has been done gives you this.

00:35:58.289 --> 00:36:05.010
Let's say higher power to unravel this hidden patterns and you care less about advanced statistics.

00:36:05.099 --> 00:36:05.550
exactly.

00:36:05.550 --> 00:36:16.650
And the other thing is, at least in my eyes, if I know how the algorithm sees the problem, I might be able to figure out any phenomena or sub phenomena that they haven't known before.

00:36:16.650 --> 00:36:18.510
And maybe this is why I'm very can methods.

00:36:18.666 --> 00:36:25.777
They're by design, very conservative because we have to be conservative, maybe we could, if we know why we don't have to be extremely conservative.

00:36:25.867 --> 00:36:38.166
And plus we know now something new that we didn't know before we can figure out why this thing, when I think of AI I always think of, I tool that can give me an answer to why this thing I did to why I didn't know this before.

00:36:38.887 --> 00:36:40.356
What's this new knowledge to me.

00:36:40.356 --> 00:36:42.483
I'm not really looking for c orrelation.

00:36:42.516 --> 00:36:58.606
I mean, my earlier work was heavily data-driven correlation because I mean, I didn't know better, but nowadays it's just makes more sense for some papers, of course data-driven would work because paper itself is for a data driven problem, the overall idea should not always be data driven.

00:36:58.606 --> 00:37:00.407
It should be more, much more than that.

00:37:00.436 --> 00:37:02.327
It should always be advanced in science.

00:37:02.356 --> 00:37:07.516
How can we advance our knowledge having to spend and thousands of dollars?

00:37:07.516 --> 00:37:12.277
And, and here's the thing, Somebody that does experiment now, years from now, it's forgotten.

00:37:12.527 --> 00:37:16.836
Somebody goes back and repeat the experiment, and they get the grant to redo the experiment again.

00:37:16.867 --> 00:37:18.637
Or they don't, you know, expand an experiment.

00:37:19.266 --> 00:37:23.617
papers wouldn't be published something in a way after a few months, it's shelved away.

00:37:23.786 --> 00:37:28.327
It's in a database sciencedirect or Springer, it's, it's being online.

00:37:28.447 --> 00:37:29.496
We rarely visit.

00:37:30.150 --> 00:37:32.820
But why do we have to continue doing the cycle all over again?

00:37:32.820 --> 00:37:39.539
If you accumulate our knowledge and we are able to come up with something new, then we can different directions that we haven't seen before.

00:37:40.384 --> 00:38:00.278
I've started with classifying this, AI into supervised unsupervised semi-supervised and you were talking about regression classification on other ways to formally classify this, but I think the true first choice is, do you use it for discovery or you do it to calculate something, you know?

00:38:00.309 --> 00:38:15.376
And, I think that's the first thing, because if you just want to figure out a number out of a very complex array of results, you have obtained that you're unable to process, in other way, because the correlations or somethingare multidimmensional.

00:38:15.721 --> 00:38:25.818
then you probably are seeking a different path than when you try to employ this method to do, to find unexpected and discover something.

00:38:25.951 --> 00:38:29.451
and, as an engineer, I would like to have better numbers.

00:38:29.451 --> 00:38:41.181
I would not necessarily be happy, discovering a completely new failure mode because that's, uh, I mean, I made the wrong, or we're kind of screwed as a humanity if I do.

00:38:41.670 --> 00:38:48.221
But as a scientist, I would like, I maybe care less about the numbers and they would care more about discovery.

00:38:48.641 --> 00:38:52.643
And, coming back to your thought about,, collectively adding to that.

00:38:53.393 --> 00:38:53.664
Okay.

00:38:53.713 --> 00:38:55.844
To what extent the data from the past.

00:38:56.398 --> 00:38:56.938
Exactly.

00:38:57.193 --> 00:39:00.943
To what extent you can take your papers from 50, 60 seventies.

00:39:00.974 --> 00:39:08.353
I don't know, from last IAFSS and use them to develop your own models is how big of an issue is that?

00:39:08.483 --> 00:39:12.923
The thing is because we talk about fire, it's, a very read it.

00:39:12.954 --> 00:39:16.164
Like it's a very niche area that, and it's a very expensive area.

00:39:16.164 --> 00:39:31.797
We don't really have a lot of experiments, or like low-risk tests that we can use, but we do have some, if you want to start with machine learning with a goal to come up, let's say with a black box surrogate that tells you failure mode or failure time, rather than doing a very lengthy calculation.

00:39:33.027 --> 00:39:39.387
You really have to do what you have, and those would be experiments that the old experiments now, the good thing is the following.

00:39:39.867 --> 00:39:45.141
The good thing is those experiments are the same ones that we use now by our care.

00:39:45.141 --> 00:39:51.291
So, you know, in a way, w we have some kind of similarity, the, on the other, on the opposite side material is different.

00:39:51.291 --> 00:39:55.130
Like for instance, concrete 50 years ago is, is really different than the concrete we have now.

00:39:55.161 --> 00:40:02.121
So the experiments 50 years ago which is also, which is what the codes are built on, are not built on your new experiments.

00:40:02.686 --> 00:40:03.266
I'm happy.

00:40:03.266 --> 00:40:03.985
You've added that.

00:40:04.201 --> 00:40:07.771
Codes are built on very, very old expert in the sixties, fifties, seventies.

00:40:08.010 --> 00:40:14.641
So even the code, why would you apply the code now when it's 60 years later, how does that accumulate to what we have now?

00:40:15.096 --> 00:40:15.456
Wow.

00:40:15.661 --> 00:40:22.068
a way,  I know it may not be as comprehensive or as accurate as doing the knowledge that we have now.

00:40:22.398 --> 00:40:40.596
However, this is the practice that we're using And if you want to compare, if you really want to compare, let's say code that a procedure against machine learning to be fair, you have to kind of use the data that develop that that a provision, which is the old data and apply to the algorithm and see, the comparison.

00:40:40.596 --> 00:40:44.666
You have to have a fair line of comparison to, this is how it works but however,.

00:40:44.735 --> 00:40:47.976
Am I happy with using 50 year old experiments?

00:40:48.096 --> 00:40:49.025
I'm not happy.

00:40:49.025 --> 00:40:50.465
No, but this is the ones we have.

00:40:50.465 --> 00:40:51.936
And this is the standard we have to use.

00:40:52.175 --> 00:40:59.088
Maybe in the future, it will be different on the good side, on the other dimension, using all the experiments.

00:40:59.213 --> 00:41:02.483
And let's say that you have two columns, one very, very old one, very, very new.

00:41:03.076 --> 00:41:05.759
The failure mode is not going to be something new.

00:41:06.509 --> 00:41:08.398
It would have still failed in the same manner.

00:41:08.789 --> 00:41:12.688
However, the time Um, it's going to be different because we have different chemicals.

00:41:12.688 --> 00:41:15.088
We have different stuff now that we use in our material.

00:41:15.266 --> 00:41:17.358
at different loading, different, temperature.

00:41:17.699 --> 00:41:21.748
However, the we're subjecting this element to is the same.

00:41:21.748 --> 00:41:23.619
They still have the same chamber compartment.

00:41:24.159 --> 00:41:27.398
So it's not really that we're completely using something different.

00:41:27.878 --> 00:41:29.469
It's just, there are some differences.

00:41:29.469 --> 00:41:35.559
And even if you want to do a, like a statistical analysis, like a meta analysis, have to compare different data from different experiments.

00:41:35.739 --> 00:41:53.646
And this is, again, this is why using data-driven analysis is a little bit itchy for me now, because I want to know why, like, at least in my mind, I want to know why this column fails so if it's fail that, years ago, or now there has the mechanism is not going to be some, some new physics.

00:41:53.746 --> 00:41:56.315
It's going to be something that maybe we haven't seen before.

00:41:56.916 --> 00:41:57.965
how can I get to that?

00:41:57.965 --> 00:42:18.427
Something that we haven't seen before by using the same old methods that we have been using for 50 years by now, we would have figured out, you know, So maybe if we use something, method, maybe we can see a little bit different and maybe that a little bit of difference would open up a new experiments for us or a new research area for us that we can apply news That's interesting.

00:42:18.427 --> 00:42:23.340
And, for, for me personally, I really liked the Xinyans,  I'm in the world of smoke control.

00:42:23.340 --> 00:42:39.574
And I really loved, how he perceived that the CFD could let to let's say more capable algorithms that would predict, the smoke behavior in a compartment, giving you a number the time to, for the layer to fall down or some tenability criteria to be breached.

00:42:40.143 --> 00:42:56.547
And for example, one of mine main, , areas of research is car parks . I engineer a lot of smoke control in car parks and our limitations is usually that we take a car park, we do 2, 3, 4, 5 CDs in it, for a certain size of the fire.

00:42:57.148 --> 00:43:01.108
And I assume if I did a sufficiently large amount of.

00:43:01.572 --> 00:43:02.552
Simulations.

00:43:02.913 --> 00:43:06.242
And then they have received a new car park with the new architecture.

00:43:06.242 --> 00:43:11.822
With, I know I've performed 1, 2, 3 simulations in that carpark.

00:43:12.182 --> 00:43:18.813
The algorithm could technically take over and tell me what would happen in like a thousand different scenarios in that car park.

00:43:18.992 --> 00:43:20.905
could you use it in like, this, like.

00:43:21.085 --> 00:43:21.206
Yeah.

00:43:21.775 --> 00:43:27.186
instance, now what we're doing, we have a database of our, say columns, 200 columns.

00:43:27.266 --> 00:43:33.452
We could come up, we can ask the algorithm to simulate a data worth of testing, 5,000 columns or 10,000.

00:43:34.338 --> 00:43:41.148
So this way, I'm trying to capture as many interactions between the features or between the parameters as I couldn't have done using experiments.

00:43:41.748 --> 00:43:47.688
But however, I still to have that baseline that at least as the algorithm, this is the main, this is the map.

00:43:47.728 --> 00:43:52.847
This is the average of the distribution of the possible that I could see before.

00:43:53.202 --> 00:43:53.532
Hm.

00:43:53.557 --> 00:43:53.927
Xinyan.

00:43:54.311 --> 00:44:02.617
Any the problem that you're going to be using simulation for it will be expensive not only expensive, you'll have to continue to do it all over and over again.

00:44:03.307 --> 00:44:09.623
And you know, once we're done with this, say with your design, you throw away this, you know, maybe you clear your desks and yeah.

00:44:09.623 --> 00:44:09.882
Yeah.

00:44:09.887 --> 00:44:10.367
throw it away.

00:44:10.547 --> 00:44:37.643
But if you have this, let's say on an annual basis, and let's say you designed 50 structures or like 50 cases, the simulation that you have is very valuable information because if you accumulate them by five or six or seven years, you'll have a very, very good database that you can teach an algorithm Maybe figured out something that we haven't seen before, maybe come up with some kind of a faster approach to solve the small problem to figure out at least what could be.

00:44:37.673 --> 00:44:42.233
And this would be interesting to me, what would be a severe case for this parking structure?

00:44:42.893 --> 00:44:51.413
it without having to house it or doing many minutes in a CFDs before, I may be wrong, but do you exactly know which one would be a severe case right now of hand?

00:44:51.463 --> 00:44:54.032
It is expert judgment and you use the design.

00:44:54.932 --> 00:45:04.503
That's the thing, because if you could use this technique to expand the number of investigated cases, you can start talking about risk and probabilities.

00:45:05.163 --> 00:45:10.742
Like a fire of this probability is giving these consequences with this confidence.

00:45:11.132 --> 00:45:16.885
And the fire of this probability is giving you these consequences at these intervals, and then you go, and it was beautiful.

00:45:16.885 --> 00:45:24.585
You could ask the AI, please test any smoke exhaust capacity from this amount of CFMs to this amount of CFMs.

00:45:24.976 --> 00:45:32.635
And, then it will tell you, okay, if you increase the ventilation twice, you decrease your probabilities by this amount.

00:45:32.635 --> 00:45:36.985
And if you increase it sevenfold, it doesn't change much from the previous case.

00:45:37.465 --> 00:45:43.076
So you, you start to get much more detailed.

00:45:43.230 --> 00:45:53.460
Outcome of your analysis then you would have from investigating multiple points, even if you are the best CFD engineer in the world, because it's not the tool that limits you.

00:45:53.880 --> 00:46:02.760
It's the capabilities of running multiple parallel cases that's essentially limiting and that there was a solution to that.

00:46:02.789 --> 00:46:04.170
There exists solutions to that.

00:46:04.590 --> 00:46:07.907
There was PhD student of Bart Merci and now Dr.

00:46:08.056 --> 00:46:13.166
Bart van Weyenberge, who was doing his PhD on  a response surface technique.

00:46:13.467 --> 00:46:21.713
It's a statistical technique where you can map, certain, uh, inputs to certain outputs of, multidimensional, uh, surfaces.

00:46:22.342 --> 00:46:24.233
And from that you can buy running.

00:46:24.713 --> 00:46:35.632
Let's say 10,CFDs or 20 CFDs you can predict the outcome of multiple CFDs, but it still requires you to solve for a certain geometry in here with machine learning.

00:46:35.632 --> 00:46:42.543
Maybe you could use results from different building to enhance your knowledge about this particular building.

00:46:42.559 --> 00:46:45.099
I mean, it's amazing because it already did.

00:46:45.099 --> 00:46:49.179
The response surface seemed like magic, and this is magic plus.

00:46:49.480 --> 00:46:52.380
It's if that happens it's gonna be amazing.

00:46:52.650 --> 00:46:53.880
And I really wish It happened.

00:46:54.869 --> 00:46:58.612
So if I wanted it, to happen What should I do now?

00:46:58.822 --> 00:47:00.081
Should I go learn coding?

00:47:00.101 --> 00:47:01.121
what's the first step?

00:47:01.152 --> 00:47:01.947
And, Let's assume.

00:47:01.947 --> 00:47:03.597
I don't know anything about coding.

00:47:03.597 --> 00:47:03.987
I don't know.

00:47:03.987 --> 00:47:04.378
Python.

00:47:04.378 --> 00:47:04.768
I don't know.

00:47:04.768 --> 00:47:06.018
R I don't know anything.

00:47:06.315 --> 00:47:07.402
but I just love this.

00:47:07.431 --> 00:47:09.742
Where, where should they, what should they do with myself?

00:47:09.891 --> 00:47:11.842
to me honest, I learned everything on YouTube.

00:47:12.081 --> 00:47:13.311
They have five Yeah.

00:47:13.791 --> 00:47:15.952
for every kind of For everything.

00:47:15.981 --> 00:47:16.192
Yeah.

00:47:16.702 --> 00:47:17.661
can either learn from them.

00:47:17.692 --> 00:47:21.891
The good news is Uh, at this moment, we don't really develop algorithms.

00:47:21.891 --> 00:47:23.632
So there are many, many codes.

00:47:23.641 --> 00:47:26.842
Like if you go to SciKit there is like the cause already there.

00:47:26.842 --> 00:47:28.831
So you can just copy paste them, Hmm.

00:47:28.851 --> 00:47:29.992
add your data, run it.

00:47:30.001 --> 00:47:32.864
And then, you know, you can fine tune, a few parameters.

00:47:32.925 --> 00:47:34.005
You should be good to go.

00:47:34.005 --> 00:47:35.594
It's not, it's something that's complex.

00:47:35.925 --> 00:47:50.954
Once you start to go maybe into explainability confidence, trust, then you have to have some kind of a very good background when it comes to math or calculus, because they're, at that point, it's not just ask them, but as applying, in other words, it's more on the development side.

00:47:51.585 --> 00:47:56.324
if you want to figure it out causality, or for instance, cause and effect, this is at least what I'm trying to do.

00:47:56.715 --> 00:48:01.414
you want to figure out cause and effect, then you have to have much more higher advancements  for coding.

00:48:01.894 --> 00:48:04.125
So the bottom box, I mean, this is what i do with my students.

00:48:04.980 --> 00:48:08.550
I'm not really expecting you to a new algorithm.

00:48:08.550 --> 00:48:09.869
If we can do that, that'd be great.

00:48:10.260 --> 00:48:14.340
However, the algorithms we have now can solve many, many, many problems.

00:48:14.369 --> 00:48:16.500
And all what you really have to do is two things.

00:48:16.500 --> 00:48:22.079
One understand how the algorithm worked its assumptions as limitation know how to apply it.

00:48:22.728 --> 00:48:29.099
You don't read it need to code it by hand because the codes are already they're available online.

00:48:29.099 --> 00:48:30.570
You can just copy paste them from there.

00:48:31.230 --> 00:48:32.760
to find your data and apply it.

00:48:32.760 --> 00:48:34.920
And then will see if you apply it.

00:48:35.340 --> 00:48:37.559
mean, I did this experiment in two of my favorites.

00:48:38.159 --> 00:48:41.820
took five or six algorithms and I applied them by default values.

00:48:41.820 --> 00:48:45.179
I just copied and pasted them on our data.

00:48:45.579 --> 00:48:46.400
And it works.

00:48:46.530 --> 00:48:59.378
I mean, you get 95%, you get 90% with very, very cheap resources, tells me that you could basically apply the same algorithm for different problems and your are gonna get very good results too.

00:48:59.831 --> 00:49:04.974
Not all the time, but at least for the most of the time, because these algorithms are extremely powerful.

00:49:05.074 --> 00:49:15.454
That's a relief in a way, you know, and I had Matt Bonner as well in here and he told the same thing that there are algorithms that exist and you can apply them.

00:49:16.054 --> 00:49:17.135
Xinyan said the same.

00:49:17.135 --> 00:49:18.925
You're the , third person to tell you the same.

00:49:18.925 --> 00:49:24.574
So I must build my brave and, and just try, I guess that's how I learned programming.

00:49:24.605 --> 00:49:29.074
Actually just, just keep trying and do as many mistakes as you can.

00:49:29.074 --> 00:49:30.724
And eventually it will work out.

00:49:31.235 --> 00:49:31.534
So.

00:49:31.670 --> 00:49:34.130
a new course next fall on machine learning.

00:49:34.849 --> 00:49:38.255
send you a link for my lecture so you can, you can attend Oh, really?

00:49:39.461 --> 00:49:40.302
That's so cool.

00:49:40.541 --> 00:49:41.742
I would appreciate that.

00:49:42.242 --> 00:49:48.032
for the end that you usually referring to resources and you have your webpage, that's very rich in resources.

00:49:48.032 --> 00:49:50.501
So I will also link to that.

00:49:50.532 --> 00:49:56.405
And, you had the paper in Fire Technology about, different types of machine learning that can be used in fire.

00:49:56.764 --> 00:50:02.885
You had this Engineer's Guide to AI and automation in construction, which was a very interesting case study.

00:50:02.885 --> 00:50:04.965
And, , it was a really nice paper.

00:50:04.994 --> 00:50:16.105
W what else should I refer the audience to, to read up on, on this, I really feel , for fire there is going to be,  the mechanistic, , review paper is a very good one , for a beginner.

00:50:16.164 --> 00:50:20.244
I know I sent out professor Rein in like a very short letter.

00:50:20.485 --> 00:50:21.925
It's going to be published very, very soon.

00:50:21.925 --> 00:50:23.875
So that, that would be a compliment to that one.

00:50:23.875 --> 00:50:25.795
Once it is, I'll send you a link for that.

00:50:26.581 --> 00:50:33.481
Engineer's Guide is, is one of my PIs, or at least when I think about it, this was highlight of 2021.

00:50:33.481 --> 00:50:35.711
For me, that, that paper Is the one really?

00:50:36.641 --> 00:50:37.452
I really liked it.

00:50:37.452 --> 00:50:37.722
There.

00:50:37.952 --> 00:50:43.621
it, even the times that I spent a lot of time on the title, because I figured that would be something very close to my heart.

00:50:44.041 --> 00:50:46.742
Uh, there is a third paper it would be mapping function.

00:50:46.742 --> 00:50:49.262
So it, I think it's after naming function.

00:50:49.831 --> 00:50:58.172
this is where we're trying to use more of a cause and effect kind of machine learning, or how can we arrive at that cause and effect having to hassle with coding.

00:50:58.172 --> 00:51:15.172
And we can actually figure out a pathways between different algorithms come up with a function or mathematical expression that can convey to us some kind of, a formula or at least can give us, because if you think of the, out of the fanatical with them, it's a number that for us engineers would like to see.

00:51:15.291 --> 00:51:16.632
We're trained in  formulas.

00:51:16.751 --> 00:51:25.105
We see that for instance, this is the format that you can apply get an output machine and it gives you a number and hence, this is the hesitation.

00:51:25.105 --> 00:51:27.565
We can see why we can see how it was.

00:51:28.092 --> 00:51:29.141
In mapping function.

00:51:29.141 --> 00:51:35.952
It's a way that it can translate algorithmic logic from a black box into a function that we can see.

00:51:36.431 --> 00:51:46.628
if you can see it, you can see the interaction between the parameters, your life had to feel much more comfortable applying a function, as opposed to applying a complete black box that we don't know why does okay.

00:51:46.628 --> 00:51:48.128
That's really, really good.

00:51:48.128 --> 00:51:53.768
And, some external or, uh, resources like maybe YouTube channel or something that you can recommend send you these.

00:51:53.768 --> 00:51:54.998
I have them on my bookmarks.

00:51:54.998 --> 00:51:58.400
I'll send you a, really you a link Fantastic.

00:51:58.400 --> 00:52:02.932
I'll put it in the show notes and I, I hope, , someone will, find it useful.

00:52:02.932 --> 00:52:07.882
And, uh, I really, I really appreciate you, you sharing this knowledge okay.

00:52:07.882 --> 00:52:09.932
Nasser, that was a great talk.

00:52:09.932 --> 00:52:18.722
And I learned something about AI today and, maybe I'm one step closer to understanding how it can be applied in my field.

00:52:18.722 --> 00:52:24.485
And I guess there's many heads buzzing now, how can this be implemented in their fields?

00:52:24.940 --> 00:52:25.360
okay.

00:52:25.481 --> 00:52:29.320
Thank you for joining us in the Fire Science Show and I hope you had a great time.

00:52:29.320 --> 00:52:31.070
I had the lot, Thank you very much.

00:52:31.070 --> 00:52:32.360
I appreciate your reaching out.

00:52:32.481 --> 00:52:33.351
I appreciate your show.

00:52:33.351 --> 00:52:33.771
Very good.

00:52:33.771 --> 00:52:36.490
I mean, I always watch watch the shows when you post them on Twitter.

00:52:36.521 --> 00:52:38.101
It's a very really?

00:52:38.190 --> 00:52:38.820
That's cool.

00:52:39.501 --> 00:52:40.130
I like that.

00:52:40.130 --> 00:52:42.291
You just do one thing.

00:52:42.380 --> 00:52:46.190
It's like different components within the fire wrodl so it's much more informative this way.

00:52:46.561 --> 00:52:46.860
Yeah.

00:52:47.121 --> 00:52:47.840
that very much.

00:52:48.036 --> 00:52:48.817
Thank you so much.

00:52:48.847 --> 00:52:49.356
Cheers, man.

00:52:49.436 --> 00:52:49.706
Bye-bye.

00:52:50.802 --> 00:52:51.432
And that's it.

00:52:51.603 --> 00:52:53.233
Well, what a discussion that was.

00:52:53.282 --> 00:52:58.193
Maybe I just should open some python right now and then start digging into that.

00:52:58.193 --> 00:53:03.262
I'm really excited about this world of fire science and the, possibilities it brings.

00:53:03.922 --> 00:53:18.378
MZ has used AI in so many different aspects of fire engineering, like literally go to his webpage and check out his papers, the variety of topics, where this method was used and considered useful.

00:53:18.643 --> 00:53:22.063
It's just amazing how wide this technology is.

00:53:22.422 --> 00:53:23.563
Of course there are caveats.

00:53:23.893 --> 00:53:26.563
You need to worry about the data quality.

00:53:26.563 --> 00:53:29.083
You need to worry about what the algorithms have not seen.

00:53:29.413 --> 00:53:32.023
I hope you've picked up these things from our discussion.

00:53:32.023 --> 00:53:37.652
That technology is powerful, but just as powerful as the algorithm and as powerful as the data that fuels it.

00:53:38.163 --> 00:53:42.702
And by far, most importantly, as powerful as the person who's using that.

00:53:43.123 --> 00:53:43.873
So if.

00:53:44.733 --> 00:53:47.682
I don't know what you're doing and you drop machine learning on that.

00:53:48.132 --> 00:53:58.097
Well, you're going to have a machine learned no idea what you're doing, but if you know what you are doing and you know what you're looking for, it's just hell, a complex problem to dig into that.

00:53:58.516 --> 00:54:04.086
Well then machine learning and artificial intelligence, maybe your best future friend.

00:54:04.679 --> 00:54:09.635
Now this talk today, I think it's a part of a mini series in the podcast.

00:54:10.175 --> 00:54:24.034
If you remember, I had an episode with Xinyan Huang from Hong Kong Polytechnic university, with whom I have discussed artificial intelligence and its potential use for smoke control and fire engineering at large.

00:54:24.454 --> 00:54:28.565
So you definitely, definitely should check that episode if you've missed that one.

00:54:29.164 --> 00:54:31.295
And I had an episode with Matt Bonner.

00:54:31.474 --> 00:54:41.985
My friend from Imperial College London, who has also used machine learning algorithms to investigate database of facade fires that we have built together.

00:54:42.434 --> 00:54:50.655
And it was also quite an interesting to see how well the artificial intelligence has carried the task that took  such a long time.

00:54:51.105 --> 00:54:54.855
So I'm really, really happy to have this in the podcast portfolio.

00:54:55.405 --> 00:54:58.465
I think these three episodes go together very well.

00:54:58.465 --> 00:55:01.315
And yeah, if you haven't heard them, absolutely.

00:55:01.315 --> 00:55:10.764
After this one, you need to tune in, into Xinyan's episode and Matt's episode, I'm going to drop the links in the show notes And yeah, that's it for today.

00:55:10.855 --> 00:55:13.284
I hope you've enjoyed it as much as I did.

00:55:13.394 --> 00:55:17.864
As usual, next episode, we'll be waiting for you here next Wednesday.

00:55:18.224 --> 00:55:21.164
Looking forward to that and yeah.

00:55:21.195 --> 00:55:21.914
See you around.

00:55:21.945 --> 00:55:22.574
Thank you for listening.

Chapters

No chapters available.