1 00:00:00,06 --> 00:00:03,03 - [Narrator] In the last video, we set up our environment, 2 00:00:03,03 --> 00:00:07,06 but now let's talk about what an AI agent actually is. 3 00:00:07,06 --> 00:00:09,07 And you can think of an agent as a program 4 00:00:09,07 --> 00:00:13,01 that understands your goals, makes a plan to achieve them, 5 00:00:13,01 --> 00:00:15,01 and then takes action. 6 00:00:15,01 --> 00:00:17,01 At the heart of the agent is this core loop. 7 00:00:17,01 --> 00:00:19,07 It consists of these four key steps, 8 00:00:19,07 --> 00:00:23,08 understanding your intent, making a plan, using tools, 9 00:00:23,08 --> 00:00:25,07 and then reflecting on what it did. 10 00:00:25,07 --> 00:00:27,04 So let's talk through these steps. 11 00:00:27,04 --> 00:00:29,07 First of all is understanding intent, 12 00:00:29,07 --> 00:00:32,03 and this all starts with you, the user. 13 00:00:32,03 --> 00:00:35,00 You're going to give the agent a task like finding a place 14 00:00:35,00 --> 00:00:36,08 for hike near where I am. 15 00:00:36,08 --> 00:00:39,07 The first thing that the agent has to do is understand 16 00:00:39,07 --> 00:00:42,00 what it is that you're asking for. 17 00:00:42,00 --> 00:00:45,01 It needs to figure out what your intent is. 18 00:00:45,01 --> 00:00:47,06 And in this case, the intent is to go for a hike 19 00:00:47,06 --> 00:00:49,04 at a nearby national park. 20 00:00:49,04 --> 00:00:51,06 Now you can use a large language model to do that, 21 00:00:51,06 --> 00:00:54,00 so there's many different ways that I could say this, 22 00:00:54,00 --> 00:00:55,06 I want to hike, I want to go for a hike, 23 00:00:55,06 --> 00:00:57,02 please help me with a hike. 24 00:00:57,02 --> 00:00:59,04 And the nice thing there about the large language model 25 00:00:59,04 --> 00:01:02,05 is that instead of doing the traditional natural language 26 00:01:02,05 --> 00:01:04,08 processing to try and figure out what you said, 27 00:01:04,08 --> 00:01:07,03 the LLM can actually artificially understand 28 00:01:07,03 --> 00:01:10,08 that on your behalf and turn that into an intent. 29 00:01:10,08 --> 00:01:13,06 So now once the agent has derived the intent 30 00:01:13,06 --> 00:01:15,03 from what you said you want to do, 31 00:01:15,03 --> 00:01:17,02 now it needs to make a plan. 32 00:01:17,02 --> 00:01:19,07 Now this plan means it breaks down the task 33 00:01:19,07 --> 00:01:22,03 into smaller manageable steps. 34 00:01:22,03 --> 00:01:23,08 So for our hiking agent, 35 00:01:23,08 --> 00:01:25,09 that plan might look something like this. 36 00:01:25,09 --> 00:01:28,01 Well, first we need to know where the user is, 37 00:01:28,01 --> 00:01:29,05 what's their location. 38 00:01:29,05 --> 00:01:32,04 Secondly, we need to get the weather at that location. 39 00:01:32,04 --> 00:01:34,01 Maybe today is not a good day for a hike 40 00:01:34,01 --> 00:01:36,06 because it's raining or snowing, that kind of thing. 41 00:01:36,06 --> 00:01:38,05 So if it is suitable, let's continue. 42 00:01:38,05 --> 00:01:41,09 And to continue would be we'll find nearby national parks, 43 00:01:41,09 --> 00:01:45,02 and we'll get more details about them and then we'll parse 44 00:01:45,02 --> 00:01:48,09 the details about those parks to get hiking suggestions, 45 00:01:48,09 --> 00:01:53,05 and then give the user some opinionated suggestions 46 00:01:53,05 --> 00:01:56,00 based on the weather, based on the parks 47 00:01:56,00 --> 00:01:57,06 and those types of things. 48 00:01:57,06 --> 00:02:00,01 Now all of that involves using tools. 49 00:02:00,01 --> 00:02:03,03 So in order for the agent to be able to make a plan, 50 00:02:03,03 --> 00:02:05,09 it needs to know what tools are available to it. 51 00:02:05,09 --> 00:02:08,03 And once it knows the tools are available to it, 52 00:02:08,03 --> 00:02:10,07 then it can as part of the planning process, 53 00:02:10,07 --> 00:02:12,03 use those tools. 54 00:02:12,03 --> 00:02:14,08 Now typically those tools are going to be APIs, 55 00:02:14,08 --> 00:02:15,06 of course, right? 56 00:02:15,06 --> 00:02:18,07 So the agent needs to know how to call an API for weather, 57 00:02:18,07 --> 00:02:20,09 how to call the API for location, 58 00:02:20,09 --> 00:02:22,07 and then even like the national parks, 59 00:02:22,07 --> 00:02:25,07 how to call the API for that and to parse the results. 60 00:02:25,07 --> 00:02:28,05 And this is where agents get really, really powerful. 61 00:02:28,05 --> 00:02:31,07 Those tools are what the agent uses to intersect, 62 00:02:31,07 --> 00:02:34,04 and interact with the outside world. 63 00:02:34,04 --> 00:02:36,00 Now, a tool, like I said, could be anything 64 00:02:36,00 --> 00:02:39,02 from a simple function like convert Celsius to Fahrenheit, 65 00:02:39,02 --> 00:02:41,02 or it could be a full blown API like searching 66 00:02:41,02 --> 00:02:42,06 for national parks. 67 00:02:42,06 --> 00:02:45,01 And our hiking agent is going to use those tools 68 00:02:45,01 --> 00:02:47,06 that I've mentioned, understand where the user is, 69 00:02:47,06 --> 00:02:49,02 fetch the weather where they are, 70 00:02:49,02 --> 00:02:51,08 and then search for national parks if appropriate, 71 00:02:51,08 --> 00:02:54,07 to be able to hike based on the current weather. 72 00:02:54,07 --> 00:02:56,05 Once it's done all of this kind of thing, 73 00:02:56,05 --> 00:02:58,05 it then needs to reflect, 74 00:02:58,05 --> 00:03:01,04 it will take look at the results of its actions and ask, 75 00:03:01,04 --> 00:03:03,06 well, did I accomplish this goal? 76 00:03:03,06 --> 00:03:06,08 LLMs are well known for hallucination, for example, 77 00:03:06,08 --> 00:03:09,03 and this is one of those steps where you can take a look 78 00:03:09,03 --> 00:03:12,01 at what has happened, look at the user's intent, 79 00:03:12,01 --> 00:03:15,03 look at the details from the tools, look at the suggestions, 80 00:03:15,03 --> 00:03:18,09 and say, hey, I really did think I accomplished the goal. 81 00:03:18,09 --> 00:03:22,00 Or if not, then go back and try again, 82 00:03:22,00 --> 00:03:24,05 or maybe ask the user for more details. 83 00:03:24,05 --> 00:03:26,02 Are you okay to hike in the snow? 84 00:03:26,02 --> 00:03:27,05 Those type of things. 85 00:03:27,05 --> 00:03:29,08 So that's where the reflection part of it 86 00:03:29,08 --> 00:03:31,02 will often give you the results 87 00:03:31,02 --> 00:03:33,09 because the results are good, but if they're not good, 88 00:03:33,09 --> 00:03:35,07 a really good agent will go back, 89 00:03:35,07 --> 00:03:38,06 take a look at its results and go back and maybe replan 90 00:03:38,06 --> 00:03:41,04 and maybe indeed as part of replanning also go back 91 00:03:41,04 --> 00:03:44,06 to maybe get extra intent from the user. 92 00:03:44,06 --> 00:03:47,06 So to recap, this is how you should be thinking 93 00:03:47,06 --> 00:03:49,07 when you're building agents to break it down 94 00:03:49,07 --> 00:03:51,02 into these four steps. 95 00:03:51,02 --> 00:03:54,03 To understand the user's intent, to plan, 96 00:03:54,03 --> 00:03:57,00 to act on that plan, and then to reflect. 97 00:03:57,00 --> 00:03:59,08 And this is the basic loop that all agents will follow. 98 00:03:59,08 --> 00:04:02,03 So next up, we're finally going to get our hands dirty 99 00:04:02,03 --> 00:04:05,00 and we'll write our very first agent.