We teach your AI to think like an expert

We find where your AI fails and work with professionals from your sector to create the data that fixes it.

If I cancel my fiber now, do I have to pay anything?

Outdated document

It read the 2024 terms. Since March 2026, leaving in the first 12 months costs up to €150.

Source: Terms of service · March 2026 · clause 9

Do I keep my number if I switch to you?

Correct

Matches the number portability page.

Source: Help center · portability

My internet is slow today. Can you speed it up?

Invented answer

It can't change a line. It should run a line test or book a technician.

Source: Assistant scope · allowed actions

Backed by

FUESCYL · Junta de Castilla y LeónIniciativa Campus EmprendedorSantander XUniversidad de Valladolid
Fundación UVaAyuntamiento de ValladolidConsolida Startup

Finding out your AI is wrong through a customer complaint means you're already too late. We catch those errors before they cost you money and claims, and build the data your system needs so they don't happen again.

How it works

We gather your documents, the questions your customers really ask and our experts' judgment, and turn them into an exam with the correct answers. We put it to your AI as a customer would, grade every answer against its source and give you a report on what fails and how to fix it.

Your documentsReal questionsOur expertsExamquestions and solutionsYour AIasked like a customerGraded answerschecked against the sourceReportscore · causes · fixesYour documentsReal questionsOur expertsExamquestions and solutionsYour AIasked like a customerGraded answerschecked against the sourceReportscore · causes · fixes

An exam built from your own documents

Every question comes with its correct answer and the exact page it comes from, so no one can dispute the result. When your AI gets something wrong, we tell you which of these five ways it failed, so you know whether to fix your documents or your system.

Five ways an AI gets it wrong

Is my jewelry covered if my home is burgled?

Your AI says

Yes, all your jewelry is covered, with no limit.

Correct answer, from your documents

Home policy · art. 14

Jewelry is covered up to €3,000, and only if it was kept in a safe.

A network of experts in your sector

We work with practicing professionals in each sector. They write and review the exam questions and, from the failures we find, create the data your AI is missing to answer correctly: from documentation that didn't exist to real cases solved step by step.

Who writes

  • Insurance adjusters
  • Claims handlers
  • Lawyers
  • Tax advisers
  • After-sales engineers
  • Doctors and nurses

What they build

  • Questions and answersPairs grounded in your documents and daily practice
  • Missing documentationWhat your documents don't say, or say twice
  • Step-by-step reasoningReal cases solved the way a professional would
  • PreferencesRanking and correcting candidate answers
  • Larger examsNew questions and edge cases every month

What you receive

A clear report written for whoever owns the assistant, not for engineers: your score, every wrong answer next to its source, why it happened and what to fix first. We walk your team through it live, and the exam is yours to keep.

The reportscore, errors and fixes
Your examevery question, yours to keep
A live walkthroughwe go through it with your team
Monthly alertswhen something breaks

Continuous improvement

Your documents change, your AI gets updated and your customers ask new things. Every month we add about 25 questions, run the exam again, alert you if anything breaks, and our experts create the data that closes each new gap.

Frequently asked questions

What does Caudals do, in plain words?

We check whether your company's AI assistant gives your customers the right answers. We write an exam from your own documents, ask your assistant the questions a real customer would ask, and grade every answer against the page it should come from. You get a report on what it gets wrong, why, and how to fix it. Then our sector experts build the data that closes those gaps.

Which AI systems can you evaluate?

Any system that answers questions in text or by voice: chatbots on your website, app or WhatsApp, phone agents, and internal assistants your team uses for policies, procedures or technical manuals. We focus on questions where a wrong answer costs money: prices, fees, coverage, eligibility, deadlines and specifications.

Do you need access to our systems or customer data?

Not to start. The free diagnostic uses only your public assistant, your public documentation and made-up customer profiles. For internal systems we work under an NDA and a data processing agreement on EU infrastructure, and delete your material when the work ends. If you prefer, we hand you the questions and you run them yourselves.

How much does it cost to start?

Nothing. The diagnostic is free: we put 40 questions to your assistant and send you the results within 48 hours. A pilot evaluation, with 150 to 300 questions reviewed by experts, takes about two weeks and has a fixed price agreed before we start. We never bill by the hour.

Will anyone else see our results?

Your results stay with you: we never publish or share them, or even say you are a customer, without your written consent. What the people who use your AI system will notice is the improvement: correct answers more often, fewer repeated questions, and more trust in your service.

Turn every mistake your AI makes into an improvement

Start with a free diagnostic: we put forty questions to your assistant and, within 48 hours, tell you what fails and what data it needs to improve.