The flagship

The part most teams skip: testing your AI.

Anyone can wire an LLM into a product in a weekend. Proving it stays accurate, safe, and on-brand under real traffic is the hard part, and it is where most teams have nothing. That is the gap I close - I build the evaluation layer around your AI.

Eval suites that score every model output against your ground truth, run in CI, and fail the build when quality drops.
Red-teaming for prompt injection, jailbreaks, and unsafe output, before someone on the internet finds them first.
RAG & retrieval evaluation - is it actually grounding answers in your data, or confidently making things up?
Regression gates so a prompt tweak or a model upgrade cannot silently break what already worked.

I already run this discipline in production - a test suite I grew from 80 to 700+ tests, with AI tooling I built myself on the Anthropic API.

  • Anthropic API
  • OpenAI API
  • Promptfoo
  • Braintrust
  • LangSmith
  • DeepEval
  • Prompt engineering

Inbox open

Let's make your next
release boring.

Mail is the fastest way to reach me. Tell me what you are shipping and where it hurts, and I will tell you how I would test it.

Email tykhonkozachenko@gmail.com

Remote · worldwide · CEST +352 661 566 607

T.