1 00:00:00,00 --> 00:00:02,40 - [Instructor] So now that we have extracted 2 2 00:00:02,40 --> 00:00:03,60 some of these episodes 3 3 00:00:03,60 --> 00:00:04,500 and the information about them, 4 4 00:00:04,500 --> 00:00:06,330 for each podcast episode 5 5 00:00:06,330 --> 00:00:08,460 we need to do three different things. 6 6 00:00:08,460 --> 00:00:10,890 We need to get the audio file, 7 7 00:00:10,890 --> 00:00:12,210 the podcast audio file, 8 8 00:00:12,210 --> 00:00:13,620 and download it. 9 9 00:00:13,620 --> 00:00:16,110 And then we'll need to transcribe that audio. 10 10 00:00:16,110 --> 00:00:17,700 And once we have a transcription, 11 11 00:00:17,700 --> 00:00:19,170 which will be a text file, 12 12 00:00:19,170 --> 00:00:21,30 we can upload that transcription 13 13 00:00:21,30 --> 00:00:22,650 to our file search store. 14 14 00:00:22,650 --> 00:00:24,810 So these transcripts of these audio files 15 15 00:00:24,810 --> 00:00:28,230 are essentially what the semantic search will happen over 16 16 00:00:28,230 --> 00:00:30,330 in order to retrieve information 17 17 00:00:30,330 --> 00:00:32,490 for a particular user query. 18 18 00:00:32,490 --> 00:00:34,500 So I will show you in just a little bit 19 19 00:00:34,500 --> 00:00:37,320 how to do this across all podcast episodes 20 20 00:00:37,320 --> 00:00:38,460 in an RSS feed, 21 21 00:00:38,460 --> 00:00:40,260 kind of in a big bulk upload. 22 22 00:00:40,260 --> 00:00:41,520 But before we get into that 23 23 00:00:41,520 --> 00:00:43,650 and having everything written out in nice functions 24 24 00:00:43,650 --> 00:00:44,700 that we can just call, 25 25 00:00:44,700 --> 00:00:46,350 I'm going to do this step by step 26 26 00:00:46,350 --> 00:00:47,573 for just a single episode. 27 27 00:00:47,573 --> 00:00:51,30 So just so we get like a handle of every single step 28 28 00:00:51,30 --> 00:00:52,350 that we need to follow 29 29 00:00:52,350 --> 00:00:53,340 and how this works, 30 30 00:00:53,340 --> 00:00:55,380 we're just going to do this for a single episode, 31 31 00:00:55,380 --> 00:00:57,480 and we'll do it for that AlphaFold episode. 32 32 00:00:57,480 --> 00:01:01,950 So I'm going to extract the first episode in that list, 33 33 00:01:01,950 --> 00:01:04,980 and we'll just call it ep, E-P. 34 34 00:01:04,980 --> 00:01:07,530 And now from that dictionary, 35 35 00:01:07,530 --> 00:01:10,170 we need to extract the audio URL. 36 36 00:01:10,170 --> 00:01:13,260 So what you'll find is that these RSS feeds 37 37 00:01:13,260 --> 00:01:15,720 have a link to the audio 38 38 00:01:15,720 --> 00:01:17,520 for every single episode. 39 39 00:01:17,520 --> 00:01:18,930 So if we type this out here, 40 40 00:01:18,930 --> 00:01:21,180 you can see there is a URL 41 41 00:01:21,180 --> 00:01:23,430 that we extracted. 42 42 00:01:23,430 --> 00:01:26,59 And if we open up this audio URL. 43 43 00:01:26,59 --> 00:01:28,170 (upbeat music) 44 44 00:01:28,170 --> 00:01:30,90 - [John] Enabled you to solve the grand challenge 45 45 00:01:30,90 --> 00:01:31,140 and that this was kind of. 46 46 00:01:31,140 --> 00:01:33,750 - [Hannah] Did you know that you need to? 47 47 00:01:33,750 --> 00:01:35,430 - [John] No, you do all those things right. 48 48 00:01:35,430 --> 00:01:37,560 - [Instructor] All right, so that's 47 minutes 49 49 00:01:37,560 --> 00:01:42,390 of the podcast audio in this URL. 50 50 00:01:42,390 --> 00:01:44,100 Okay, and so we're going to go ahead 51 51 00:01:44,100 --> 00:01:47,100 and just create a name for our audio file. 52 52 00:01:47,100 --> 00:01:48,960 We're going to call it episode_0, 53 53 00:01:48,960 --> 00:01:51,120 and then for our transcription as well. 54 54 00:01:51,120 --> 00:01:53,160 So we'll use these in just a minute. 55 55 00:01:53,160 --> 00:01:55,770 So now that we have acquired the audio URL, 56 56 00:01:55,770 --> 00:01:56,790 the next thing we need to do 57 57 00:01:56,790 --> 00:01:59,70 is actually download that audio. 58 58 00:01:59,70 --> 00:02:01,890 So I'm going to do this with the requests library 59 59 00:02:01,890 --> 00:02:04,860 and I'm going to just download this information in chunks, 60 60 00:02:04,860 --> 00:02:08,370 and I'm going to save it as episode_0.mp3. 61 61 00:02:08,370 --> 00:02:10,350 So you can see on the left side over here, 62 62 00:02:10,350 --> 00:02:12,360 our audio file has now been downloaded, 63 63 00:02:12,360 --> 00:02:16,20 and it's downloaded locally to my Colab environment. 64 64 00:02:16,20 --> 00:02:17,550 And so the next thing we need to do 65 65 00:02:17,550 --> 00:02:20,460 is we need to actually transcribe this audio 66 66 00:02:20,460 --> 00:02:22,80 because, again, we want text. 67 67 00:02:22,80 --> 00:02:25,230 And that is what the model will be able to search over 68 68 00:02:25,230 --> 00:02:26,670 when a user asks a query. 69 69 00:02:26,670 --> 00:02:29,790 So to do this, I'm going to read in the audio bites 70 70 00:02:29,790 --> 00:02:31,980 of my MP3 file, 71 71 00:02:31,980 --> 00:02:36,420 and then I'm going to use Gemini 3 Pro Preview again, 72 72 00:02:36,420 --> 00:02:39,30 but to actually create the transcription. 73 73 00:02:39,30 --> 00:02:41,190 So I'm calling generate_content. 74 74 00:02:41,190 --> 00:02:43,50 Again, that is the method 75 75 00:02:43,50 --> 00:02:46,170 of text generation in the Gemini API. 76 76 00:02:46,170 --> 00:02:48,900 And we pass in Gemini 3 Pro Preview. 77 77 00:02:48,900 --> 00:02:51,570 If you want to experiment with another model, 78 78 00:02:51,570 --> 00:02:53,40 you can definitely try that out 79 79 00:02:53,40 --> 00:02:55,260 and just change the model's string name. 80 80 00:02:55,260 --> 00:02:56,760 And then I'm going to instruct the model 81 81 00:02:56,760 --> 00:02:59,340 to, "Generate a transcript of this audio file. 82 82 00:02:59,340 --> 00:03:01,560 Make sure to transcribe the entire file. 83 83 00:03:01,560 --> 00:03:02,610 Do not repeat yourself. 84 84 00:03:02,610 --> 00:03:05,310 And label the speakers." 85 85 00:03:05,310 --> 00:03:07,890 And so after we pass in that prompt, 86 86 00:03:07,890 --> 00:03:10,200 we also need to pass in the audio file, 87 87 00:03:10,200 --> 00:03:11,610 which is right here. 88 88 00:03:11,610 --> 00:03:13,650 So this is going to take about a minute or so. 89 89 00:03:13,650 --> 00:03:16,500 It's a long audio file that needs to be processed 90 90 00:03:16,500 --> 00:03:19,80 and also needs to be transcribed. 91 91 00:03:19,80 --> 00:03:20,70 So while that happens, 92 92 00:03:20,70 --> 00:03:21,360 I'm going to quickly jump over 93 93 00:03:21,360 --> 00:03:23,190 and show you this GitHub repo 94 94 00:03:23,190 --> 00:03:25,980 that my awesome teammate Mark McDonald made. 95 95 00:03:25,980 --> 00:03:28,380 This has all of the code that I'm showing you today, 96 96 00:03:28,380 --> 00:03:31,830 but written a little bit more production friendly. 97 97 00:03:31,830 --> 00:03:33,750 But you can see that there's some information here. 98 98 00:03:33,750 --> 00:03:35,310 And if we scroll down here, 99 99 00:03:35,310 --> 00:03:37,890 you can see that there are basically two steps. 100 100 00:03:37,890 --> 00:03:39,510 You can call ingest.py, 101 101 00:03:39,510 --> 00:03:41,940 and you just pass in your RSS feed. 102 102 00:03:41,940 --> 00:03:43,380 So you don't need to go through all the steps 103 103 00:03:43,380 --> 00:03:44,970 that I'm walking you through right now, 104 104 00:03:44,970 --> 00:03:47,40 you can just pass in the URL as is, 105 105 00:03:47,40 --> 00:03:49,170 and then you can call query.py 106 106 00:03:49,170 --> 00:03:50,400 and ask a question. 107 107 00:03:50,400 --> 00:03:53,40 So I definitely encourage you to take a look at this. 108 108 00:03:53,40 --> 00:03:53,873 If you jump into here, 109 109 00:03:53,873 --> 00:03:55,590 you'll see basically all of the same code 110 110 00:03:55,590 --> 00:03:56,423 I'm running through, 111 111 00:03:56,423 --> 00:03:58,380 but it just looks a lot cleaner 112 112 00:03:58,380 --> 00:04:00,210 and might be a little bit harder to follow. 113 113 00:04:00,210 --> 00:04:01,380 So that's why we're running through it 114 114 00:04:01,380 --> 00:04:03,510 step by step over here. 115 115 00:04:03,510 --> 00:04:05,880 So looks like this is still running. 116 116 00:04:05,880 --> 00:04:08,70 Let's let this keep running. 117 117 00:04:08,70 --> 00:04:12,210 And while that transcription chugs away, 118 118 00:04:12,210 --> 00:04:14,190 we can scroll ahead a little bit 119 119 00:04:14,190 --> 00:04:16,350 and take a look at the next few steps. 120 120 00:04:16,350 --> 00:04:17,760 So what we're going to do next 121 121 00:04:17,760 --> 00:04:20,10 with this transcription file once we have it 122 122 00:04:20,10 --> 00:04:23,700 is we will upload it to our file search store. 123 123 00:04:23,700 --> 00:04:26,190 Oh, and actually, looks like I stalled for long enough 124 124 00:04:26,190 --> 00:04:27,570 and the transcription is done. 125 125 00:04:27,570 --> 00:04:29,400 So let's go ahead and take a look. 126 126 00:04:29,400 --> 00:04:32,160 So we'll call response.text. 127 127 00:04:32,160 --> 00:04:35,160 And you can see we have a transcription. 128 128 00:04:35,160 --> 00:04:37,20 So first we have John Jumper, 129 129 00:04:37,20 --> 00:04:38,977 and he's talking about, 130 130 00:04:38,977 --> 00:04:40,140 "I think we'll get this ability 131 131 00:04:40,140 --> 00:04:42,210 to poke the cell in exciting ways, 132 132 00:04:42,210 --> 00:04:43,920 to interrogate it, et cetera. 133 133 00:04:43,920 --> 00:04:45,540 We have Hannah Fry, 134 134 00:04:45,540 --> 00:04:47,820 who is the host of "The Google DeepMind" podcast, 135 135 00:04:47,820 --> 00:04:49,147 and you can see she's saying, 136 136 00:04:49,147 --> 00:04:51,750 "Welcome to 'The Google DeepMind' podcast," 137 137 00:04:51,750 --> 00:04:52,770 et cetera, et cetera. 138 138 00:04:52,770 --> 00:04:56,520 If we click here, we can see the whole transcription. 139 139 00:04:56,520 --> 00:04:58,290 So now that we've done that, 140 140 00:04:58,290 --> 00:05:02,400 we are going to create a transcription file 141 141 00:05:02,400 --> 00:05:04,740 and upload this to the file search store. 142 142 00:05:04,740 --> 00:05:10,440 So I am first going to write this transcription. 143 143 00:05:10,440 --> 00:05:12,00 So you can see right over here on the left, 144 144 00:05:12,00 --> 00:05:15,750 we just had this file transcript_0.txt open. 145 145 00:05:15,750 --> 00:05:17,160 So let's take a look. 146 146 00:05:17,160 --> 00:05:20,250 So I just added some nice metadata at the top. 147 147 00:05:20,250 --> 00:05:22,440 I added the title of the podcast episode, 148 148 00:05:22,440 --> 00:05:23,370 the name of the podcast, 149 149 00:05:23,370 --> 00:05:24,210 and then the date. 150 150 00:05:24,210 --> 00:05:25,920 And then now we have our transcript. 151 151 00:05:25,920 --> 00:05:28,110 So you can see John and Hannah 152 152 00:05:28,110 --> 00:05:30,690 having a back-and-forth conversation. 153 153 00:05:30,690 --> 00:05:33,120 So that is what I was doing right here in this block here. 154 154 00:05:33,120 --> 00:05:35,970 I was just adding this extra information, 155 155 00:05:35,970 --> 00:05:37,890 the title, the podcast and the date, 156 156 00:05:37,890 --> 00:05:39,660 in addition to that transcription 157 157 00:05:39,660 --> 00:05:42,300 that we got from Gemini. 158 158 00:05:42,300 --> 00:05:44,70 And then the next thing I'm going to do 159 159 00:05:44,70 --> 00:05:46,50 is get some extra metadata 160 160 00:05:46,50 --> 00:05:49,770 because when we upload this to the file search store, 161 161 00:05:49,770 --> 00:05:51,810 we want to provide some extra information, 162 162 00:05:51,810 --> 00:05:54,960 like the title of the episode name, the podcast, 163 163 00:05:54,960 --> 00:05:56,370 just as metadata. 164 164 00:05:56,370 --> 00:05:58,530 We didn't do this in the bake shop example, 165 165 00:05:58,530 --> 00:05:59,700 but you can see right here 166 166 00:05:59,700 --> 00:06:01,710 that I'm passing in this config here 167 167 00:06:01,710 --> 00:06:05,460 with some extra information about this text file. 168 168 00:06:05,460 --> 00:06:06,990 And so once we've done that, 169 169 00:06:06,990 --> 00:06:09,00 we have our transcript uploaded 170 170 00:06:09,00 --> 00:06:10,890 to our file search store, 171 171 00:06:10,890 --> 00:06:13,470 which means that we can now ask a question 172 172 00:06:13,470 --> 00:06:16,950 about the information in this podcast episode. 173 173 00:06:16,950 --> 00:06:19,740 So again, I'm calling generate_content, 174 174 00:06:19,740 --> 00:06:23,790 and I'm going to call this on the Model Gemini 3 Pro Preview. 175 175 00:06:23,790 --> 00:06:24,817 And my prompt here is, 176 176 00:06:24,817 --> 00:06:27,450 "What did John and Hannah talk about?" 177 177 00:06:27,450 --> 00:06:30,390 And hopefully since we have passed in our file search store 178 178 00:06:30,390 --> 00:06:31,230 right here, 179 179 00:06:31,230 --> 00:06:34,350 we should get a response that relates to AlphaFold. 180 180 00:06:34,350 --> 00:06:37,770 So let's see. 181 181 00:06:37,770 --> 00:06:38,737 So you can see it says, 182 182 00:06:38,737 --> 00:06:40,680 "Based on the AlphaFold grand challenge 183 183 00:06:40,680 --> 00:06:42,810 to Nobel Prize with John Jumper," blah, blah, blah. 184 184 00:06:42,810 --> 00:06:45,510 We have a summary kind of what the two of them talked about. 185 185 00:06:45,510 --> 00:06:46,380 And just so we know, 186 186 00:06:46,380 --> 00:06:47,550 this actually did work 187 187 00:06:47,550 --> 00:06:51,90 if we were to delete this part of our request 188 188 00:06:51,90 --> 00:06:54,210 that includes the file search tool, 189 189 00:06:54,210 --> 00:06:56,310 as well as the name of our store. 190 190 00:06:56,310 --> 00:06:57,780 And if we call this again, 191 191 00:06:57,780 --> 00:07:00,210 probably the model is going to come back with a response 192 192 00:07:00,210 --> 00:07:02,430 that is like, "Who are John and Hannah?" 193 193 00:07:02,430 --> 00:07:04,290 So it won't have access anymore 194 194 00:07:04,290 --> 00:07:05,940 to the grounded information 195 195 00:07:05,940 --> 00:07:08,250 from our podcast transcription. 196 196 00:07:08,250 --> 00:07:09,690 So let's take a look and see 197 197 00:07:09,690 --> 00:07:11,220 what the model thinks or who it guesses 198 198 00:07:11,220 --> 00:07:13,800 John and Hannah maybe are. 199 199 00:07:13,800 --> 00:07:16,290 All right, so here we have a response from the model, 200 200 00:07:16,290 --> 00:07:19,230 and it says, "Because John and Hannah are very common names, 201 201 00:07:19,230 --> 00:07:20,430 I need a little more context." 202 202 00:07:20,430 --> 00:07:23,640 So you can see if we don't pass in the file search store, 203 203 00:07:23,640 --> 00:07:25,500 the model has no idea which John and Hannah 204 204 00:07:25,500 --> 00:07:27,150 we are actually talking about. 205 205 00:07:27,150 --> 00:07:29,310 But if we have that data in our search store, 206 206 00:07:29,310 --> 00:07:32,220 it can access the correct information 207 207 00:07:32,220 --> 00:07:34,00 and get the context we're looking for.