From the blog
What we're learning, in the open.
Field notes on agent readiness, evaluation craft, and the mechanisms that make evidence trustworthy, from the team building Stratix.
Stratix · Method
RAG Evaluation Basics: The Metrics That Matter
Retrieval quality, answer quality, and the gap between them. The handful of measurements that actually tell you whether a RAG pipeline is ready.
Read the post
Stratix · Method
LLM Testing: Unit Tests, Evals, and the Gap Between Them
Stratix · Evaluation
Your Agent Eval Can't Tell If a Change Helped or Hurt
Stratix · Evaluation
What Are AI Evals? A Plain-English Guide
Stratix · Evaluation
LLM Evaluation for Beginners: How It Actually Works
Stratix · Method
Your RAG Pipeline Scores 0.91 Faithfulness. It Still Hallucinates.
Stratix · Evaluation
The Score on the Box: Why Benchmark Numbers Stop Meaning What You Think
Stratix · Method
AI Agent Testing Breaks the Moment Agents Remember
Stratix · Method
LLM as a Judge: Two Runs, Two Scores, No Answer
Stratix · Product
Stratix Ships Compass: Model Selection That Weighs What Your Industry Needs
Stratix · Evaluation
"AI evaluation platform" Now Describes Three Different Jobs
Stratix · Evaluation
Every Major Provider Shipped Multi-Agent. None Shipped Evaluation for It.
Stratix · Cup
Stratix Cup Season 1: Six Rounds of LLM Self-Improvement in Public
No posts in this topic yet. More coming.