1 00:00:00,06 --> 00:00:02,09 - [Instructor] Let's talk about agent intelligence, 2 00:00:02,09 --> 00:00:05,00 how to make your agent smarter. 3 00:00:05,00 --> 00:00:08,04 In simple terms, this comes down to which model you choose 4 00:00:08,04 --> 00:00:10,07 and how you led that reason. 5 00:00:10,07 --> 00:00:13,05 A large language model is the core brain of your agent, 6 00:00:13,05 --> 00:00:15,05 which orchestrates the interaction flow 7 00:00:15,05 --> 00:00:17,07 and handles the tool calls. 8 00:00:17,07 --> 00:00:22,06 In our n8n setup, the agent was powered by Gemini 2.5 Flash, 9 00:00:22,06 --> 00:00:26,00 which is a modern, lightweight and fast model. 10 00:00:26,00 --> 00:00:27,07 But there's a large spectrum. 11 00:00:27,07 --> 00:00:31,09 There are larger models like Gemini 2.5 Pro, GPT 5, 12 00:00:31,09 --> 00:00:36,02 or Claude 4.5 Sonnet that have far more parameters. 13 00:00:36,02 --> 00:00:38,07 Choosing a bigger model is your first lever 14 00:00:38,07 --> 00:00:40,07 for more intelligence. 15 00:00:40,07 --> 00:00:42,06 Bigger models generally tend to handle 16 00:00:42,06 --> 00:00:44,07 nuance and context better, 17 00:00:44,07 --> 00:00:47,04 but they're slower and more expensive. 18 00:00:47,04 --> 00:00:49,04 Smaller ones respond instantly, 19 00:00:49,04 --> 00:00:52,07 but they may miss detail or make simpler assumptions. 20 00:00:52,07 --> 00:00:54,06 The key is finding the right balance 21 00:00:54,06 --> 00:00:56,08 that fits your use case. 22 00:00:56,08 --> 00:00:58,05 The good thing is that in n8n, 23 00:00:58,05 --> 00:01:01,02 you can try out different models easily. 24 00:01:01,02 --> 00:01:03,05 Just switch the model node in your agent, 25 00:01:03,05 --> 00:01:05,04 and you're all set. 26 00:01:05,04 --> 00:01:08,04 Another lever you have is to adjust the reasoning time 27 00:01:08,04 --> 00:01:10,03 if the model allows this. 28 00:01:10,03 --> 00:01:11,06 This controls essentially 29 00:01:11,06 --> 00:01:14,03 how long the model is allowed to think. 30 00:01:14,03 --> 00:01:16,01 You can imagine it like giving the model 31 00:01:16,01 --> 00:01:17,09 more mental scratch space. 32 00:01:17,09 --> 00:01:21,02 It generates more tokens to explore different avenues 33 00:01:21,02 --> 00:01:22,05 to solve a given problem 34 00:01:22,05 --> 00:01:25,04 before arriving at the final answer. 35 00:01:25,04 --> 00:01:28,00 More reasoning steps usually improve accuracy, 36 00:01:28,00 --> 00:01:30,04 but at the cost of latency. 37 00:01:30,04 --> 00:01:32,07 Less reasoning gives instant results, 38 00:01:32,07 --> 00:01:35,02 but sometimes only at surface level. 39 00:01:35,02 --> 00:01:37,09 Here's what that looks like inside n8n. 40 00:01:37,09 --> 00:01:41,00 You can pick a model that supports different reasoning modes 41 00:01:41,00 --> 00:01:43,06 like OpenAI's GPT 5 models. 42 00:01:43,06 --> 00:01:45,03 In this case, you can choose between 43 00:01:45,03 --> 00:01:48,00 low, medium, or high reasoning efforts, 44 00:01:48,00 --> 00:01:50,02 depending on whether you need quick replies 45 00:01:50,02 --> 00:01:52,05 or more reliable reasoning. 46 00:01:52,05 --> 00:01:55,02 Note that n8n sometimes lags a little behind 47 00:01:55,02 --> 00:01:56,02 in the customization 48 00:01:56,02 --> 00:02:00,01 compared to the original model provider APIs. 49 00:02:00,01 --> 00:02:01,07 Besides picking a different model 50 00:02:01,07 --> 00:02:03,08 or adjusting the reasoning time, 51 00:02:03,08 --> 00:02:07,02 advanced teams can also try to fine-tune a model. 52 00:02:07,02 --> 00:02:10,01 That means training the model on your own examples, 53 00:02:10,01 --> 00:02:13,06 internal tools, interaction flows, or tone. 54 00:02:13,06 --> 00:02:16,06 This is very powerful, but also much more complex. 55 00:02:16,06 --> 00:02:17,09 Usually a later step 56 00:02:17,09 --> 00:02:20,04 once your base agent works reliably 57 00:02:20,04 --> 00:02:23,03 and you have collected data to fine-tune on. 58 00:02:23,03 --> 00:02:25,02 In the end it's about trade-offs. 59 00:02:25,02 --> 00:02:28,02 You can increase intelligence by choosing a stronger model 60 00:02:28,02 --> 00:02:30,01 or give it more time to think, 61 00:02:30,01 --> 00:02:31,07 but every game needs more time 62 00:02:31,07 --> 00:02:34,01 and creates higher inference costs. 63 00:02:34,01 --> 00:02:35,09 The art is to find the sweet spot 64 00:02:35,09 --> 00:02:37,06 where your agent is accurate enough 65 00:02:37,06 --> 00:02:40,03 and fast enough to feel responsive, 66 00:02:40,03 --> 00:02:43,09 and you'll need a few good iterations to figure this out. 67 00:02:43,09 --> 00:02:45,07 My tip, start simple. 68 00:02:45,07 --> 00:02:49,02 Use a small fast model like Gemini 2.5 Flash 69 00:02:49,02 --> 00:02:50,06 to get things running, 70 00:02:50,06 --> 00:02:53,08 and once your setup is stable and you can measure results, 71 00:02:53,08 --> 00:02:56,04 try larger models or longer reasoning time 72 00:02:56,04 --> 00:02:58,03 to see if accuracy improves. 73 00:02:58,03 --> 00:03:01,05 And don't forget to experiment with different providers too. 74 00:03:01,05 --> 00:03:05,06 Models like Claude 4.5 Haiku or GPT-5 mini 75 00:03:05,06 --> 00:03:07,08 sometimes follow your custom instructions 76 00:03:07,08 --> 00:03:10,09 or your specific tools better out of the box. 77 00:03:10,09 --> 00:03:12,03 So test and compare. 78 00:03:12,03 --> 00:03:14,03 That's how you'll find your sweet spot 79 00:03:14,03 --> 00:03:18,00 between speed, cost, and reliability.