{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "## Evaluating CrewAI's `crew` (end-to-end)\n", "\n", "In this notebook we will demonstrate how you can run evaluations on crews using datasets from Confident AI and DeepEval's dataset iterator." ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Install dependencies:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "!pip install -U deepeval -U crewai ipywidgets --quiet" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Set your OpenAI API key:" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "import os\n", "\n", "os.environ[\"OPENAI_API_KEY\"] = \"\"" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Create a crew:\n", "\n", "This is a simple crew with a single agent and a single task." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from crewai import Task, Crew, Agent\n", "\n", "agent = Agent(\n", " role=\"Consultant\",\n", " goal=\"Write clear, concise explanation.\",\n", " backstory=\"An expert consultant with a keen eye for software trends.\",\n", ")\n", "\n", "task = Task(\n", " description=\"Explain the given topic: {topic}\",\n", " expected_output=\"A clear and concise explanation.\",\n", " agent=agent,\n", ")\n", "\n", "crew = Crew(agents=[agent], tasks=[task])\n", "\n", "result = crew.kickoff(\n", " inputs={\"topic\": \"What is the biggest open source database?\"}\n", ")\n", "print(result)" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Evaluate the agent\n", "\n", "To evaluate CrewAI's `crew`:\n", "\n", "1. Instrument the application (using `from deepeval.integrations.crewai import instrument_crewai`)\n", "2. Supply metrics to `kickoff`.\n", "\n", "\n", "> (Pro Tip) View your Agent's trace and publish test runs on [Confident AI](https://www.confident-ai.com/). Apart from this you get an in-house dataset editor and more advaced tools to monitor and enventually improve your Agent's performance. Get your API key from [here](https://app.confident-ai.com/)\n" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "os.environ[\"CONFIDENT_API_KEY\"] = \"\"" ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from deepeval.integrations.crewai import instrument_crewai\n", "\n", "instrument_crewai()" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Using a dataset from Confident AI:\n", "\n", "For demo purposes, we will use a public dataset from Confident AI. You can use your own dataset as well. Refer to the [docs](https://deepeval.com/docs/evaluation-end-to-end-llm-evals#setup-your-test-environment) to learn more about how to create your own dataset." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from deepeval.dataset import EvaluationDataset\n", "\n", "dataset = EvaluationDataset()\n", "dataset.pull(alias=\"topic_agent_queries\", public=True)" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### Run evaluations:\n", "\n", "We will use the `AnswerRelevancyMetric` to evaluate the crew. Dataset iterator will yield golden examples from the dataset." ] }, { "cell_type": "code", "execution_count": null, "metadata": {}, "outputs": [], "source": [ "from deepeval.metrics import AnswerRelevancyMetric\n", "\n", "for golden in dataset.evals_iterator():\n", " result = crew.kickoff(\n", " inputs={\"topic\": golden.input}, metrics=[AnswerRelevancyMetric()]\n", " )" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "Congratulation! You have just evaluated your first CrewAI's `crew` using Deepeval. Try changing Hyperparameters, Agents, Tasks, Metrics and see how your agent performs." ] } ], "metadata": { "kernelspec": { "display_name": ".venv", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 3 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3", "version": "3.12.10" } }, "nbformat": 4, "nbformat_minor": 2 }