1 00:00:00,06 --> 00:00:03,00 - [Instructor] Now, let's look at how to observe your agent, 2 00:00:03,00 --> 00:00:06,08 what it's doing when it succeeds and when it silently fails. 3 00:00:06,08 --> 00:00:09,02 Monitoring and observability are what turn 4 00:00:09,02 --> 00:00:12,03 an experimental workflow into something reliable 5 00:00:12,03 --> 00:00:14,09 you can run in production. 6 00:00:14,09 --> 00:00:18,06 If you can't see what your agent does, you can't improve it. 7 00:00:18,06 --> 00:00:20,09 Let's set some terminology straight. 8 00:00:20,09 --> 00:00:22,09 Monitoring shows what happens, 9 00:00:22,09 --> 00:00:25,04 the runs, the errors, the timing. 10 00:00:25,04 --> 00:00:28,01 Observability explains why it happens, 11 00:00:28,01 --> 00:00:31,08 connecting those runs to the agent's reasoning and tool use. 12 00:00:31,08 --> 00:00:34,00 And finally, evaluation checks 13 00:00:34,00 --> 00:00:35,09 whether the outcome is correct. 14 00:00:35,09 --> 00:00:38,00 Together these give you the full picture, 15 00:00:38,00 --> 00:00:41,08 what happened, why, and if it worked as intended. 16 00:00:41,08 --> 00:00:43,07 First, let's talk about monitoring. 17 00:00:43,07 --> 00:00:46,09 You'll monitor your agent in two stages, 18 00:00:46,09 --> 00:00:48,07 at once while building it, 19 00:00:48,07 --> 00:00:52,00 catching tool call issues and logic errors early, 20 00:00:52,00 --> 00:00:55,07 and then later in production, making sure it stays stable 21 00:00:55,07 --> 00:00:58,08 and behaves predictably with real users. 22 00:00:58,08 --> 00:01:01,01 Now, let's look at monitoring and development. 23 00:01:01,01 --> 00:01:05,02 Here, monitoring usually means opening n8n's execution logs 24 00:01:05,02 --> 00:01:06,09 and reading what happened. 25 00:01:06,09 --> 00:01:08,08 Did the agent called the right tool? 26 00:01:08,08 --> 00:01:10,05 Did it sent the right payload? 27 00:01:10,05 --> 00:01:11,09 You're looking for patterns 28 00:01:11,09 --> 00:01:13,07 where it misunderstands instructions 29 00:01:13,07 --> 00:01:16,02 or passes wrong parameters. 30 00:01:16,02 --> 00:01:19,02 And once your agent is live, you move to metrics. 31 00:01:19,02 --> 00:01:23,01 How often does it fail? Which tools cause slowdowns? 32 00:01:23,01 --> 00:01:25,08 You can hook n8n into monitoring dashboards, 33 00:01:25,08 --> 00:01:27,08 send logs to a central store, 34 00:01:27,08 --> 00:01:31,03 or add error alerts through Slack or email. 35 00:01:31,03 --> 00:01:33,06 Now, let's talk about observability. 36 00:01:33,06 --> 00:01:37,01 Here's a great example of observability in action. 37 00:01:37,01 --> 00:01:39,07 You can see the agent's full workflow chain, 38 00:01:39,07 --> 00:01:41,09 the chat message, the tool it called, 39 00:01:41,09 --> 00:01:44,04 and the exact error it hit. 40 00:01:44,04 --> 00:01:47,01 In this case, the calculator tool failed 41 00:01:47,01 --> 00:01:51,05 because the parameter size_sqft was sent as text, 42 00:01:51,05 --> 00:01:54,07 100 square feet, instead of a plain number. 43 00:01:54,07 --> 00:01:57,04 And that's what observability gives you, 44 00:01:57,04 --> 00:02:00,09 insight into the agent's inner loop, how it used the tool, 45 00:02:00,09 --> 00:02:04,01 what inputs it passed, and where things went wrong. 46 00:02:04,01 --> 00:02:07,03 With that visibility, you can fix the issue immediately 47 00:02:07,03 --> 00:02:10,06 or add validation to prevent it in production. 48 00:02:10,06 --> 00:02:12,03 Looking at evaluation, 49 00:02:12,03 --> 00:02:14,07 evaluation is where you test your workflow 50 00:02:14,07 --> 00:02:16,04 under controlled conditions 51 00:02:16,04 --> 00:02:18,04 before or after deployment. 52 00:02:18,04 --> 00:02:21,02 You want to see whether it produces the right results, 53 00:02:21,02 --> 00:02:22,08 and this usually involves running 54 00:02:22,08 --> 00:02:25,07 a test dataset through your agent. 55 00:02:25,07 --> 00:02:28,04 Each test case includes a sample input 56 00:02:28,04 --> 00:02:30,06 and usually the expected output. 57 00:02:30,06 --> 00:02:31,06 That's how you measure 58 00:02:31,06 --> 00:02:35,00 consistency and accuracy automatically. 59 00:02:35,00 --> 00:02:38,05 Evaluations catch regressions before your users do. 60 00:02:38,05 --> 00:02:42,06 They help you compare model versions or prompts objectively, 61 00:02:42,06 --> 00:02:45,02 and they give you a measurable sense of improvement, 62 00:02:45,02 --> 00:02:47,04 not just a gut feeling. 63 00:02:47,04 --> 00:02:48,06 In the case of n8n, 64 00:02:48,06 --> 00:02:51,08 it even includes a built-in evaluations feature. 65 00:02:51,08 --> 00:02:55,02 This lets you define test cases right inside your workflow, 66 00:02:55,02 --> 00:02:59,08 specify expected outputs, and verify results after each run. 67 00:02:59,08 --> 00:03:01,08 You can explore that feature if you click 68 00:03:01,08 --> 00:03:06,01 on the Evaluations tab at the top of your workflow canvas. 69 00:03:06,01 --> 00:03:08,01 Evaluation is really what separates 70 00:03:08,01 --> 00:03:09,08 the flaky proof of concept 71 00:03:09,08 --> 00:03:12,00 from a stable production workflow. 72 00:03:12,00 --> 00:03:15,03 Without a tight monitoring and evaluation loop, 73 00:03:15,03 --> 00:03:17,07 your agent will never leave the lab. 74 00:03:17,07 --> 00:03:19,05 So make sure you set the basics right 75 00:03:19,05 --> 00:03:21,00 for observability first 76 00:03:21,00 --> 00:03:22,08 so you can monitor, evaluate, 77 00:03:22,08 --> 00:03:25,00 and then iterate when it matters.