1 00:00:00,510 --> 00:00:01,350 - [Instructor] So with that, 2 2 00:00:01,350 --> 00:00:03,870 I'm going to jump over here into Colab 3 3 00:00:03,870 --> 00:00:07,770 and we're going to get started building with the API. 4 4 00:00:07,770 --> 00:00:11,370 So the use case for today is building a podcast search. 5 5 00:00:11,370 --> 00:00:14,310 So I personally listen to a ton of podcasts. 6 6 00:00:14,310 --> 00:00:15,780 It's very unstructured data. 7 7 00:00:15,780 --> 00:00:18,150 I'm often listening while I'm like doing the dishes 8 8 00:00:18,150 --> 00:00:19,440 or driving my car, 9 9 00:00:19,440 --> 00:00:20,820 and so I might hear something, 10 10 00:00:20,820 --> 00:00:22,500 I won't get a chance to write it down 11 11 00:00:22,500 --> 00:00:25,410 and I always want to refer back to something that I learned. 12 12 00:00:25,410 --> 00:00:29,460 So today's use case will be taking a bunch of podcasts 13 13 00:00:29,460 --> 00:00:31,620 and making a searchable database 14 14 00:00:31,620 --> 00:00:34,950 so we can retrieve information about what we learned. 15 15 00:00:34,950 --> 00:00:37,520 To do this, we're going to be using the File Search tool 16 16 00:00:37,520 --> 00:00:39,30 in the Gemini API. 17 17 00:00:39,30 --> 00:00:41,820 This is a built-in tool that imports, chunks, 18 18 00:00:41,820 --> 00:00:43,350 and indexes your data 19 19 00:00:43,350 --> 00:00:45,930 so that that information can be retrieved 20 20 00:00:45,930 --> 00:00:47,730 based on a user's prompt. 21 21 00:00:47,730 --> 00:00:51,90 And then it's provided as context for the model prompt, 22 22 00:00:51,90 --> 00:00:53,430 which allows the model to basically give more accurate 23 23 00:00:53,430 --> 00:00:54,720 and relevant answers. 24 24 00:00:54,720 --> 00:00:55,680 So if you're thinking, 25 25 00:00:55,680 --> 00:00:58,650 well, this sounds a whole lot like RAG, you are correct. 26 26 00:00:58,650 --> 00:01:02,70 The File Search tool is essentially the managed RAG tool 27 27 00:01:02,70 --> 00:01:03,900 within the Gemini API. 28 28 00:01:03,900 --> 00:01:05,430 So that's what we're going to build today. 29 29 00:01:05,430 --> 00:01:07,260 You will need to set up your API key. 30 30 00:01:07,260 --> 00:01:09,330 I'm going to show you how to do this in Colab 31 31 00:01:09,330 --> 00:01:11,190 just in case that's the environment you're using. 32 32 00:01:11,190 --> 00:01:12,600 But if you're using something else, 33 33 00:01:12,600 --> 00:01:15,720 you can set the API key as an environment variable. 34 34 00:01:15,720 --> 00:01:17,940 In Colab, we do this as a Colab secret. 35 35 00:01:17,940 --> 00:01:20,760 So you just create a new secret here 36 36 00:01:20,760 --> 00:01:23,190 and you can call it Gemini API key 37 37 00:01:23,190 --> 00:01:26,100 and paste in the value of your key. 38 38 00:01:26,100 --> 00:01:28,260 And then we can run the cell here, 39 39 00:01:28,260 --> 00:01:31,500 which imports user data from Google Colab, 40 40 00:01:31,500 --> 00:01:34,260 and that allows us to extract our API key. 41 41 00:01:34,260 --> 00:01:36,810 So again, this is just specific to Colab. 42 42 00:01:36,810 --> 00:01:38,100 If you're doing this elsewhere, 43 43 00:01:38,100 --> 00:01:39,960 just make sure you set your API key 44 44 00:01:39,960 --> 00:01:41,730 as an environment variable. 45 45 00:01:41,730 --> 00:01:42,960 But once we've done that, 46 46 00:01:42,960 --> 00:01:46,770 we can then import the Google Gen AI SDK, 47 47 00:01:46,770 --> 00:01:50,580 and we can go ahead and create our client 48 48 00:01:50,580 --> 00:01:52,290 and pass in the API key. 49 49 00:01:52,290 --> 00:01:54,300 And again, don't hard-code this. 50 50 00:01:54,300 --> 00:01:56,280 If you're doing this in another environment, 51 51 00:01:56,280 --> 00:01:59,790 you'll just want to pass this as an environment variable. 52 52 00:01:59,790 --> 00:02:01,410 But now we have our client set up 53 53 00:02:01,410 --> 00:02:04,380 and we are ready to get started. 54 54 00:02:04,380 --> 00:02:08,970 So before we go into the more complicated podcast example, 55 55 00:02:08,970 --> 00:02:11,580 we're just going to start off using the File Search tool 56 56 00:02:11,580 --> 00:02:12,780 with a single document 57 57 00:02:12,780 --> 00:02:15,900 just to get some familiarity with how this tool works. 58 58 00:02:15,900 --> 00:02:18,600 So in order to use the File Search tool, 59 59 00:02:18,600 --> 00:02:20,520 we will go through three steps. 60 60 00:02:20,520 --> 00:02:21,353 The first thing we'll do 61 61 00:02:21,353 --> 00:02:24,30 is we will create a File Search store 62 62 00:02:24,30 --> 00:02:27,60 and then we will upload a file to the File Search store 63 63 00:02:27,60 --> 00:02:29,190 and then we can ask a question. 64 64 00:02:29,190 --> 00:02:30,270 So I'm going to get started 65 65 00:02:30,270 --> 00:02:32,760 by importing a couple libraries that I'll need. 66 66 00:02:32,760 --> 00:02:34,260 And then the next thing we'll do 67 67 00:02:34,260 --> 00:02:36,810 is create our File Search store. 68 68 00:02:36,810 --> 00:02:39,570 So the search store is essentially the database 69 69 00:02:39,570 --> 00:02:42,300 that has all of the information from our files. 70 70 00:02:42,300 --> 00:02:44,910 So it's our persistent container for these embeddings 71 71 00:02:44,910 --> 00:02:48,720 that the semantic search will operate on at inference time. 72 72 00:02:48,720 --> 00:02:51,30 So to create this, we are going to call 73 73 00:02:51,30 --> 00:02:54,990 file_search_stores.create on our client. 74 74 00:02:54,990 --> 00:02:58,50 And remember, we set up our client earlier up here, 75 75 00:02:58,50 --> 00:03:01,470 that was our Google Gen AI client up here. 76 76 00:03:01,470 --> 00:03:05,340 And we will need to give our search store a name. 77 77 00:03:05,340 --> 00:03:07,170 I'm just calling it my example store, 78 78 00:03:07,170 --> 00:03:09,960 but feel free to try out anything else you like. 79 79 00:03:09,960 --> 00:03:11,70 And then the next thing we'll do 80 80 00:03:11,70 --> 00:03:12,900 is we need to upload a file. 81 81 00:03:12,900 --> 00:03:17,490 So I need to get a file uploaded into my Colab runtime. 82 82 00:03:17,490 --> 00:03:19,680 I'm just going to drag this file 83 83 00:03:19,680 --> 00:03:22,890 and upload it right here to my Colab environment. 84 84 00:03:22,890 --> 00:03:24,900 And once I've done that, 85 85 00:03:24,900 --> 00:03:27,390 we can take a look at this file really quickly. 86 86 00:03:27,390 --> 00:03:29,160 This just has information 87 87 00:03:29,160 --> 00:03:32,460 about different baked goods at a bakery. 88 88 00:03:32,460 --> 00:03:34,710 You can see the name of the item 89 89 00:03:34,710 --> 00:03:37,410 and then the description, ingredients, and the macros. 90 90 00:03:37,410 --> 00:03:39,930 So let's jump back into Colab. 91 91 00:03:39,930 --> 00:03:41,790 Now that I've got this file uploaded, 92 92 00:03:41,790 --> 00:03:44,970 I can upload the file to my File Search store. 93 93 00:03:44,970 --> 00:03:48,632 So that is calling upload_to_file_search_store. 94 94 00:03:48,632 --> 00:03:50,490 And you'll notice that I am passing in 95 95 00:03:50,490 --> 00:03:53,160 not only the name of the file we just uploaded, 96 96 00:03:53,160 --> 00:03:55,860 that was our bake-shop.txt file, 97 97 00:03:55,860 --> 00:04:00,300 but also the specific File Search store that we created, 98 98 00:04:00,300 --> 00:04:01,920 and that's through store.name. 99 99 00:04:01,920 --> 00:04:05,340 So let's take a look at what this looks like. 100 100 00:04:05,340 --> 00:04:06,930 So if we call store.name, 101 101 00:04:06,930 --> 00:04:10,230 we'll see that there is a unique identifier 102 102 00:04:10,230 --> 00:04:13,860 created for our File Search store, 103 103 00:04:13,860 --> 00:04:15,330 and that is right here. 104 104 00:04:15,330 --> 00:04:17,130 And so this is just the identifier 105 105 00:04:17,130 --> 00:04:20,460 so that our client knows which of our File Search stores 106 106 00:04:20,460 --> 00:04:22,200 to upload this file to, 107 107 00:04:22,200 --> 00:04:24,480 because you might have multiple different ones depending on, 108 108 00:04:24,480 --> 00:04:27,270 you know, all the different use cases you're working on. 109 109 00:04:27,270 --> 00:04:30,570 So now that we've done that and our file has been uploaded, 110 110 00:04:30,570 --> 00:04:33,150 we can go ahead and ask the model a question. 111 111 00:04:33,150 --> 00:04:35,220 So I have a prompt here, which is: 112 112 00:04:35,220 --> 00:04:38,250 what item can I eat on the menu if I'm gluten-free? 113 113 00:04:38,250 --> 00:04:41,850 And in the next cell, we are going to pass that prompt 114 114 00:04:41,850 --> 00:04:44,10 to the model right here. 115 115 00:04:44,10 --> 00:04:47,190 And we'll do this with the generate_content method. 116 116 00:04:47,190 --> 00:04:48,840 So if you've used the API before, 117 117 00:04:48,840 --> 00:04:50,490 you might know that generate_content 118 118 00:04:50,490 --> 00:04:53,37 is our main text generation method. 119 119 00:04:53,37 --> 00:04:56,250 And so we can pass in the name of the model we're using, 120 120 00:04:56,250 --> 00:04:58,680 that is Gemini 3 Pro preview. 121 121 00:04:58,680 --> 00:05:01,440 And let's go ahead and execute the cell. 122 122 00:05:01,440 --> 00:05:02,273 Just to note, 123 123 00:05:02,273 --> 00:05:04,200 in addition to passing the prompt and the model, 124 124 00:05:04,200 --> 00:05:07,230 I have also passed in the File Search tool. 125 125 00:05:07,230 --> 00:05:10,950 So to do that, we create this GenerateContentConfig 126 126 00:05:10,950 --> 00:05:12,210 and we pass in tools. 127 127 00:05:12,210 --> 00:05:15,00 And the tool we're passing in is the File Search tool. 128 128 00:05:15,00 --> 00:05:16,350 And then you'll notice again, 129 129 00:05:16,350 --> 00:05:18,570 we pass in the specific name of our store. 130 130 00:05:18,570 --> 00:05:20,550 Just in case you have multiple stores 131 131 00:05:20,550 --> 00:05:22,80 associated with this client, 132 132 00:05:22,80 --> 00:05:25,170 you just need to specify the exact store. 133 133 00:05:25,170 --> 00:05:26,700 And so now that we've done this, 134 134 00:05:26,700 --> 00:05:29,190 we can take a look at the response. 135 135 00:05:29,190 --> 00:05:30,990 So we'll say response.text, 136 136 00:05:30,990 --> 00:05:32,850 and you'll see that the model says, 137 137 00:05:32,850 --> 00:05:35,700 based on the bake shop menu provided, 138 138 00:05:35,700 --> 00:05:38,130 there's one item explicitly labeled as gluten-free, 139 139 00:05:38,130 --> 00:05:38,963 blah, blah, blah. 140 140 00:05:38,963 --> 00:05:42,930 So it gives us an answer to the query that we asked. 141 141 00:05:42,930 --> 00:05:45,480 And just to make this extra clear, 142 142 00:05:45,480 --> 00:05:47,460 this whole block of text right here 143 143 00:05:47,460 --> 00:05:48,780 that I've highlighted in blue, 144 144 00:05:48,780 --> 00:05:51,300 this is the configuration of the tool. 145 145 00:05:51,300 --> 00:05:54,480 So if we were to just call the model like this 146 146 00:05:54,480 --> 00:05:56,460 without the tool in there, 147 147 00:05:56,460 --> 00:05:59,580 it wouldn't have access to this bake shop file. 148 148 00:05:59,580 --> 00:06:01,530 So if we said, what item can I eat on the menu 149 149 00:06:01,530 --> 00:06:02,550 if I'm gluten-free, 150 150 00:06:02,550 --> 00:06:04,650 it wouldn't know what we're talking about. 151 151 00:06:04,650 --> 00:06:06,330 Something else we can do 152 152 00:06:06,330 --> 00:06:10,260 is we can take a look at the metadata if we want to. 153 153 00:06:10,260 --> 00:06:13,350 So the response includes not just an answer to our question, 154 154 00:06:13,350 --> 00:06:17,610 but also citations that specify which part of the documents 155 155 00:06:17,610 --> 00:06:19,500 were used to actually generate the answer. 156 156 00:06:19,500 --> 00:06:21,120 So if you look at the grounding metadata, 157 157 00:06:21,120 --> 00:06:23,400 you'll see these chunks of data 158 158 00:06:23,400 --> 00:06:26,520 which show in the text file what was being used. 159 159 00:06:26,520 --> 00:06:29,310 And you'll also be able to see the actual text 160 160 00:06:29,310 --> 00:06:30,900 that was generated by the model 161 161 00:06:30,900 --> 00:06:33,270 and the response like this right here 162 162 00:06:33,270 --> 00:06:36,540 that is associated with those grounded chunks. 163 163 00:06:36,540 --> 00:06:39,00 So that gives us an example of how to do this 164 164 00:06:39,00 --> 00:06:40,260 with one single document. 165 165 00:06:40,260 --> 00:06:43,80 But of course, we could have just uploaded this document 166 166 00:06:43,80 --> 00:06:44,670 directly into the model's prompt. 167 167 00:06:44,670 --> 00:06:46,680 It probably would've been a whole lot faster. 168 168 00:06:46,680 --> 00:06:48,90 We didn't need to go through all the effort 169 169 00:06:48,90 --> 00:06:51,60 of actually creating this File Search store 170 170 00:06:51,60 --> 00:06:52,590 just for one single document. 171 171 00:06:52,590 --> 00:06:55,530 But creating the File Search store is really useful 172 172 00:06:55,530 --> 00:06:57,360 when we have tons and tons of documents 173 173 00:06:57,360 --> 00:07:00,00 and we don't know where the answer is exactly. 174 174 00:07:00,00 --> 00:07:02,820 And so that is a perfect kind of setup 175 175 00:07:02,820 --> 00:07:05,00 for our podcast search example.