LMTY
/
Evals
/
How we evaluate

A quality bar for AI work.

We define what good looks like for the job, then test whether the system can reach it repeatedly.

Task fit

Did the system perform the job that was actually requested?

A detailed answer to the wrong question still fails.

Context acquisition

Did it find and use the information the method needs?

Strong reasoning cannot rescue missing context forever.

Factual grounding

Are important claims supported by evidence that can be inspected?

Freshness and coverage

Is the evidence current enough for the claim being made?

Does the output make coverage limits visible instead of treating a recent check as complete knowledge?

Fact and inference

Can the reader tell what was observed and what the system concluded from it?

Reasoning quality

Does the analysis reflect the actual method for the job, or could the same generic prose have been produced for almost anything?

Prioritization

Did the system surface what matters?

A dump of everything found is not the same thing as judgment.

Usability

Can the intended person use the output in the real moment that created the job?

Missing-data behavior

When the information is weak, does the system say so?

The right failure can be more useful than invented certainty.

Action quality

Are the next steps proportionate to the evidence and the company context available?

The context ladder

We run the same work with different levels of context to see where the quality changes.

That can include user-provided material, current web research, LMTY market context, and relevant company sources.

Adversarial cases

We intentionally test conditions that make AI work look better than it is.

Stale sources. Conflicting evidence. Weak coverage. Marketing claims that sound like facts. One anecdote that looks like a trend. Company assumptions contradicted by customer evidence.

Human review

The final question is practical.

Would a strong practitioner trust this enough to use it?