Using Braintrust to evaluate agentic setups from large-scale Hugging Face data
Follow into
Save into
We pulled 1,781 real agent traces from Hugging Face into Braintrust, scored every run, and found that the wrapper around your model explains 7× more variation in success than the model itself.