Task fit
Did the system perform the job that was actually requested?
A detailed answer to the wrong question still fails.
Context acquisition
Did it find and use the information the method needs?
Strong reasoning cannot rescue missing context forever.
Factual grounding
Are important claims supported by evidence that can be inspected?
Freshness and coverage
Is the evidence current enough for the claim being made?
Does the output make coverage limits visible instead of treating a recent check as complete knowledge?
Fact and inference
Can the reader tell what was observed and what the system concluded from it?
Reasoning quality
Does the analysis reflect the actual method for the job, or could the same generic prose have been produced for almost anything?
Prioritization
Did the system surface what matters?
A dump of everything found is not the same thing as judgment.
Usability
Can the intended person use the output in the real moment that created the job?
Missing-data behavior
When the information is weak, does the system say so?
The right failure can be more useful than invented certainty.
Action quality
Are the next steps proportionate to the evidence and the company context available?
The context ladder
We run the same work with different levels of context to see where the quality changes.
That can include user-provided material, current web research, LMTY market context, and relevant company sources.
Adversarial cases
We intentionally test conditions that make AI work look better than it is.
Stale sources. Conflicting evidence. Weak coverage. Marketing claims that sound like facts. One anecdote that looks like a trend. Company assumptions contradicted by customer evidence.
Human review
The final question is practical.
Would a strong practitioner trust this enough to use it?