Learn more

For more information or if you need help retrieving your data, please contact Weights & Biases Customer Support at support@wandb.com

AGENTS

Observability tools for building autonomous AI systems that perform multi-step tasks

W&B Weave provides the tools you need to evaluate, monitor, and iterate on agentic AI systems during both development and production. With traces, scorers, guardrails, and a registry, you can confidently build and manage safe, high-quality agents and agentic workflows.

Get started

Evaluations

Agentic systems are complex and involve many different components, making eval metrics more difficult to build and track. Weave evaluates agents quickly using pre-built, third-party, or homegrown scorers, increasing iteration velocity.

Learn about Evaluations

Debugger

Agents organize their tasks into sequences of steps, each performing a variety of actions such as calling tools, reflecting on outputs, and retrieving relevant data. Debugging and iterating on these complex rollouts can be challenging with traditional call-stack views. Weave clearly visualizes complex agent rollouts to accelerate the iteration process.

Learn about Traces

Guardrails

Since LLMs are non-deterministic, you need the ability to modify agent inputs and outputs when harmful, inappropriate, or off-brand content is detected. Weave enables real-time adjustments to agent behavior to mitigate the impact of hallucinations and prompt attacks.

Learn about Guardrails

Integrations

New agent frameworks are coming to market rapidly, making it difficult to keep up. Through integrations with popular frameworks such as CrewAI and OpenAI Agents SKD, Weave helps you stay future-proof, ensuring your agents won’t become obsolete as new frameworks emerge.

See our integrations

Agent protocols

Tracing interactions between agents and tools they use can be messy. Weave cuts through the chaos—auto-logging Model Context Protocol (MCP) agent traces with one line of code. And dropping soon: plug-and-play tracing support for Google’s Agent2Agent (A2A) protocol, so you can spend less time instrumenting and more time building.

Learn how our MCP integration works

Governance

For compliance and audits, developers need the ability to rebuild specific agent versions and configurations and reproduce events arising in production. Weights & Biases acts as a system of record and allows AI developers to reproduce any task in the agent lifecycle by providing code, dataset, and metadata versioning and lineage tracking.

Learn more

Explore Weights & Biases

Learn more about Weave

  • Traces: Debug agents and AI applications
  • Evaluations: Rigorous evaluations of agentic AI systems
  • Playground: Explore prompts and models
  • Agents: Observability tools for agentic systems
  • Guardrails: Block prompt attacks and harmful outputs
  • Monitors: Continuously improve in prod
  • Models: Build and manage AI models
  • Inference: Explore hosted, open-source LLMs
  • Training: Fine-tune with serverless RL
  • Core: Optimize AI workloads

The Weights & Biases platform helps you streamline your workflow from end to end

Models

Experiments

Experiments

Track and visualize your ML experiments

Sweeps

Sweeps

Optimize your hyperparameters

Registry

Registry

Publish and share your ML models and datasets

Automations

Automations

Trigger workflows automatically

Weave

Traces

Traces

Explore and debug LLMs

Evaluations

Evaluations

Rigorous evaluations of GenAI applications

Core

Artifacts

Artifacts

Version and manage your ML pipelines

Tables

Tables

Visualize and explore your ML data

Reports

Reports

Document and share your ML insights

SDK

SDK

Log ML experiments and artifacts at scale

Get started with Agents

Get started