We test the work.
AI can produce something polished and still get the job wrong.
LMTY treats quality as something to evaluate instead of something to infer from how convincing the output sounds.
The job gets a quality bar first.
Every Skill should know what success looks like before a model is asked to perform it.
The details change by job, but the shared questions are familiar.
Did it perform the right task?
Did it use the context the method requires?
Are important claims grounded?
Is the evidence current enough?
Did it separate observation from inference?
Did it prioritize what matters?
Did it degrade honestly when information was missing?
Could the intended person actually use the result?
Context should make a measurable difference.
The same job can be tested with different levels of context.
User-provided material.
Current web research.
Persistent LMTY market context.
LMTY plus relevant company sources.
That lets us measure where additional context actually improves the work.
Failure cases matter.
We test Skills against stale information, conflicting sources, weak coverage, misleading competitor claims, anecdotes that look like trends, and missing company context.
The goal is to find the ways a plausible answer can still be wrong.
Practitioner judgment matters too.
One question sits above a lot of the detail.
Would a strong practitioner circulate this as-is?
Public scorecards come after enough testing.
We will publish scores when the fixtures, grading method, and run history are strong enough for the number to mean something.