Testing whether language model harnesses transfer the wrong strategy
Follow into
Save into

An RLM harness can carry a problem-solving strategy from one task to another. This eval tests what happens when the strategy it carries is wrong.
Follow into
Save into

An RLM harness can carry a problem-solving strategy from one task to another. This eval tests what happens when the strategy it carries is wrong.