1 00:00:00,360 --> 00:00:01,740 - [Instructor] So we're done with the bake shop. 2 2 00:00:01,740 --> 00:00:03,60 We're going to start moving into something 3 3 00:00:03,60 --> 00:00:04,350 a little bit more complicated. 4 4 00:00:04,350 --> 00:00:07,650 And again, this is going to be building our podcast search tool 5 5 00:00:07,650 --> 00:00:10,680 where we will upload a bunch of episodes of a podcast, 6 6 00:00:10,680 --> 00:00:12,570 and we will be able to ask questions 7 7 00:00:12,570 --> 00:00:15,420 about the information stored in all those podcasts. 8 8 00:00:15,420 --> 00:00:18,630 So to do that, you will need this feed parser library. 9 9 00:00:18,630 --> 00:00:20,730 So I'm going to go ahead and install this. 10 10 00:00:20,730 --> 00:00:22,290 This is just a Python library 11 11 00:00:22,290 --> 00:00:24,540 that helps do some parsing of XML files. 12 12 00:00:24,540 --> 00:00:26,460 It's just going to make our life a little bit easier 13 13 00:00:26,460 --> 00:00:28,530 when we are processing some data. 14 14 00:00:28,530 --> 00:00:30,870 So you'll want to make sure that you install that. 15 15 00:00:30,870 --> 00:00:32,850 And then we can import it. 16 16 00:00:32,850 --> 00:00:36,90 And now we will create a totally new file search store. 17 17 00:00:36,90 --> 00:00:38,460 So if you recall, we did do this earlier 18 18 00:00:38,460 --> 00:00:41,760 when we created our search store for the bake shop file. 19 19 00:00:41,760 --> 00:00:43,170 But now we're going to create a new one, 20 20 00:00:43,170 --> 00:00:46,20 and I've called this LinkedIn Live Demo store. 21 21 00:00:46,20 --> 00:00:48,510 And once that is created, we can go ahead 22 22 00:00:48,510 --> 00:00:51,360 and start downloading our podcast information. 23 23 00:00:51,360 --> 00:00:54,450 So you will need the RSS URL 24 24 00:00:54,450 --> 00:00:56,340 for the podcast that you want to use. 25 25 00:00:56,340 --> 00:00:59,370 And I'm going to show you what this one looks like. 26 26 00:00:59,370 --> 00:01:01,290 Okay, so this is the RSS feed 27 27 00:01:01,290 --> 00:01:03,600 for the Google DeepMind podcast. 28 28 00:01:03,600 --> 00:01:05,730 You can actually see the information right here. 29 29 00:01:05,730 --> 00:01:07,800 So every podcast is going to have one of these. 30 30 00:01:07,800 --> 00:01:09,180 It's just a big XML file 31 31 00:01:09,180 --> 00:01:12,00 that contains all the information about the podcast. 32 32 00:01:12,00 --> 00:01:14,490 There's tons of info here. We'll look at this more in depth. 33 33 00:01:14,490 --> 00:01:16,920 So when you do this, or if you choose to do this on your own 34 34 00:01:16,920 --> 00:01:19,800 and there's a podcast you want to be able to process, 35 35 00:01:19,800 --> 00:01:21,510 just do a Google search for, you know, 36 36 00:01:21,510 --> 00:01:23,850 the podcast name and RSS URL. 37 37 00:01:23,850 --> 00:01:26,70 There's tons of websites that will help you extract those. 38 38 00:01:26,70 --> 00:01:27,600 So that's exactly what I did for this one. 39 39 00:01:27,600 --> 00:01:31,620 I just searched for it, and it's pretty easy to find. 40 40 00:01:31,620 --> 00:01:33,960 Okay, so I'm going to execute this cell right now, 41 41 00:01:33,960 --> 00:01:36,390 and you can see I've got the printout message here, 42 42 00:01:36,390 --> 00:01:38,490 which is the title of the podcast, 43 43 00:01:38,490 --> 00:01:40,620 which was Google DeepMind: The Podcast. 44 44 00:01:40,620 --> 00:01:43,290 And so what we've done here is I've used feed parser 45 45 00:01:43,290 --> 00:01:47,40 to essentially grab the information from that XML file 46 46 00:01:47,40 --> 00:01:48,870 that I was just showing you. 47 47 00:01:48,870 --> 00:01:51,180 I only grabbed the first five entries. 48 48 00:01:51,180 --> 00:01:52,63 You can see I put a limit of five here, 49 49 00:01:52,63 --> 00:01:54,600 and that's just 'cause this is a demo; 50 50 00:01:54,600 --> 00:01:56,220 I didn't want to grab all of the data, 51 51 00:01:56,220 --> 00:01:58,320 but of course you will want to actually process 52 52 00:01:58,320 --> 00:01:59,580 all of the podcasts in there, 53 53 00:01:59,580 --> 00:02:01,920 but I just wanted to keep it simple for right now. 54 54 00:02:01,920 --> 00:02:04,260 So to understand this a little better, 55 55 00:02:04,260 --> 00:02:06,870 let's take a look at what we have done in this code cell. 56 56 00:02:06,870 --> 00:02:10,860 So I created this list here called Episodes. 57 57 00:02:10,860 --> 00:02:13,590 And we can take a look at the length here. 58 58 00:02:13,590 --> 00:02:16,680 It's five, and that makes sense because I said "limit five." 59 59 00:02:16,680 --> 00:02:19,140 So essentially I grabbed the first five episodes 60 60 00:02:19,140 --> 00:02:20,850 from that XML file. 61 61 00:02:20,850 --> 00:02:22,830 So let's go ahead and actually take a look 62 62 00:02:22,830 --> 00:02:24,990 at the first entry. 63 63 00:02:24,990 --> 00:02:28,650 And so you can see that the feed parser library 64 64 00:02:28,650 --> 00:02:30,960 very nicely took this URL here 65 65 00:02:30,960 --> 00:02:34,410 and processed everything into these dictionaries. 66 66 00:02:34,410 --> 00:02:37,200 And this is information about the first episode 67 67 00:02:37,200 --> 00:02:39,300 at the top of that URL. 68 68 00:02:39,300 --> 00:02:43,650 And the title of this podcast was all about AlphaFold. 69 69 00:02:43,650 --> 00:02:46,530 And so you can see information about the podcast: 70 70 00:02:46,530 --> 00:02:48,240 there's like a summary, 71 71 00:02:48,240 --> 00:02:50,760 there's information on when it was published, et cetera. 72 72 00:02:50,760 --> 00:02:54,00 If we did this for the next item in our list, 73 73 00:02:54,00 --> 00:02:57,00 you can see that now we have, let's see, 74 74 00:02:57,00 --> 00:02:59,340 a different podcast episode. 75 75 00:02:59,340 --> 00:03:01,410 This one is about Waymo. 76 76 00:03:01,410 --> 00:03:04,320 And so each of these items in our episode list 77 77 00:03:04,320 --> 00:03:05,430 is going to be a dictionary 78 78 00:03:05,430 --> 00:03:09,00 representing a particular episode.