We study how reasoning systems behave under pressure, then build the ones we're willing to stand behind. Nous Labs works at the intersection of formal evaluation and applied machine cognition.
Each group operates independently but reports into a shared evaluation standard, so findings in one area sharpen the others.
Long-horizon planning and multi-step inference, tested against tasks designed to break shallow pattern-matching.
Independent stress-testing of model behavior, with published methodology so results can be reproduced, not taken on faith.
Translating lab-grade reasoning capability into narrow, high-stakes decision support for regulated industries.
The evaluation harnesses, data pipelines, and compute scheduling that make rigorous research repeatable at scale.
A system that cannot explain why it reached an answer has not reasoned. It has guessed convincingly.— Nous Labs, Research Charter
We publish negative results alongside positive ones. A method that only wins in the paper isn't ready.
We earn the right to generalize by first being right in a domain where the stakes are real.
Every deployed system is evaluated by a team that did not build it, on data they did not choose.
Tell us what you're working on. We reply from a person, not a form-handler, usually within two working days.