🎉 Top 5 AI Agent for Law Firms | Read the article →
Clerx
All Posts

10/1/2026

How to Test an AI Receptionist for a Law Firm Before Go-Live

A law firm should not test an AI receptionist with one perfect call. It should test the real exceptions that determine whether intake is safe, accurate, and useful when the office is busy.

Law Firm TechnologyImplementationAI ReceptionistLegal IntakeQuality Assurance

By Attorney Michael Brunman, Co-Founder and CEO of Clerx


A law-firm AI receptionist go-live test is a controlled review of the system's answers, questions, routing, scheduling, integrations, records, privacy behavior, and exception handling before real callers depend on it. The test should cover normal inquiries and difficult edge cases, assign pass-fail criteria, identify an owner for every defect, and repeat critical scenarios after changes.


The goal is not to prove that the system can complete a scripted call. The goal is to determine whether it behaves according to the firm's rules when callers interrupt, change subjects, speak another language, request legal advice, refuse information, or need a human.

Start with a written scope

Define exactly what the receptionist is allowed to do. A Clerx AI receptionist may be configured to answer inbound calls, identify caller type, collect approved intake information, qualify against firm rules, schedule consultations, send approved links, transfer calls, and create summaries or records.


The test plan should also state what it must not do. Typical boundaries include:


  • no legal advice;
  • no prediction of outcomes;
  • no final conflicts decision;
  • no promise of representation;
  • no invented policy, fee, availability, or deadline;
  • no disclosure of client information;
  • no unsupported claim about the firm's services;
  • no bypass of a required human review.


Use the AI receptionist security checklist for vendor, access, retention, incident, and integration questions that sit outside conversation testing.

Build a test matrix, not a loose call list

A strong matrix varies four dimensions:


  1. caller type: prospect, current client, former client, referral, opposing party, vendor, spam;
  2. channel and condition: phone, chat, SMS, business hours, after hours, background noise, disconnection;
  3. matter path: qualified, unqualified, uncertain, urgent indicator, conflicts pause, payment required;
  4. outcome: booked, transferred, message taken, attorney review, declined path, opted out, safely ended.


This prevents the team from testing only the path it expects to work. The law-firm intake checklist can supply the workflow components, while the test matrix supplies the scenarios.

Define pass-fail criteria before calling

Each scenario should have observable requirements. For example, a Spanish-speaking family-law prospect might pass only if the system:


  • recognizes or confirms the preferred language;
  • identifies the caller as a new inquiry;
  • collects the approved names and county;
  • avoids legal advice;
  • recognizes a stated hearing date as an escalation trigger;
  • alerts the correct person;
  • does not book before required review;
  • creates a complete record;
  • sends no sensitive details in an exposed message.


“The call sounded good” is not a test result. Use a scorecard for accuracy, completion, tone, boundary compliance, routing, record quality, and next-step clarity.

Test category 1: identity and caller classification

Begin with the first branching decision. Test:


  • a new prospect;
  • a current client calling from a different number;
  • a family member calling for someone else;
  • a former client with a new issue;
  • an attorney referral;
  • an opposing party;
  • a vendor or court caller;
  • a person who refuses to give a name.


The system should not force every caller into new-lead intake. The new-lead and existing-client routing framework shows why misclassification creates privacy, service, and duplicate-record problems.

Test category 2: approved knowledge and legal boundaries

Ask routine operational questions about hours, locations, practice areas, consultation fees, and what to bring. Confirm that answers match the firm's current approved source.


Then ask questions the system should not answer:


  • “Do I have a case?”
  • “What should I plead?”
  • “Will this trust avoid probate?”
  • “Am I eligible for asylum?”
  • “Can my spouse take the children?”
  • “What is my deadline?”


A passing response should state the boundary clearly and route the caller to the approved next step. It should not merely add a disclaimer before giving substantive legal advice.


Test outdated and missing information as well. If the knowledge source does not contain a policy, the system should say it needs the firm to confirm rather than invent an answer.

Test category 3: intake questions and conversation control

Use callers who:


  • provide information out of order;
  • interrupt frequently;
  • change the matter type mid-call;
  • answer indirectly;
  • spell a difficult name;
  • correct an earlier answer;
  • refuse a question;
  • ask why information is needed;
  • tell a long story before giving basic facts;
  • return after a disconnection.


Check whether the workflow captures the corrected value, avoids repeating completed questions, and recovers without sounding rigid. The firm's intake form strategy should guide what is asked at each stage.

Test category 4: conflicts and sensitive information

Use common names, aliases, organizations, and multiple related parties. Test an opposing party who tries to provide a detailed narrative. Confirm that the workflow gathers only the firm's minimum preliminary data and pauses at the required gate.


Review where transcripts, recordings, summaries, and names appear. Ensure notifications do not expose sensitive facts to an unnecessarily broad group. The system may structure the record; lawyers and authorized staff retain conflicts analysis and matter acceptance.

Test category 5: qualification and disqualification

Build one scenario for every firm-approved criterion and exception. Examples include geography, matter type, case stage, consultation policy, language availability, or a service the firm does not offer.


Check both false positives and false negatives. A false positive sends an unsuitable prospect to an attorney's calendar. A false negative turns away a potentially valuable inquiry. When the facts are unclear, the safest outcome may be attorney review rather than an automated conclusion.


The AI receptionist versus intake specialist guide helps define which judgment-heavy exceptions need people.

Test category 6: transfers and escalation

Test a successful transfer and every failure mode:


  • primary recipient answers;
  • primary recipient does not answer;
  • backup answers;
  • nobody answers;
  • caller hangs up while waiting;
  • transfer occurs after hours;
  • a current-client escalation enters the wrong line;
  • the destination number is invalid;
  • the attorney asks for context before accepting.


The caller should never fall into a dead end. The law-firm intake SLA should define recipients, backups, response targets, and documentation.

Test category 7: calendar and paid consultations

For each calendar, test:


  • correct attorney and appointment type;
  • time zone;
  • minimum notice and scheduling horizon;
  • buffers and blocked time;
  • rescheduling and cancellation;
  • no availability;
  • duplicate booking attempt;
  • language-specific availability;
  • appointment requiring approval;
  • payment success, failure, and abandonment.


Use the paid-consultation workflow to verify the relationship between eligibility, payment, confirmation, and handoff. Then test the no-show reduction process for reminders and preparation instructions.

Test category 8: channel continuity

If the firm also uses website chat or text messaging, test a person moving between channels. A caller may request a link by SMS and later resume through chat. The history should remain connected where the configuration supports it.


Test consent and opt-out behavior for messaging. Confirm that the first text identifies the firm without exposing sensitive practice-area details. A reply of STOP should follow the approved suppression path.

Test category 9: integration records

The conversation can sound excellent while the downstream record fails. For each scenario, inspect the connected system.


The Clerx integrations directory lists supported practice-management, CRM, calendar, payment, automation, and messaging systems. Exact behavior depends on configuration. Confirm:


  • the correct contact or lead was found or created;
  • fields are mapped correctly;
  • corrected values replace earlier values where intended;
  • duplicate contacts are avoided;
  • summary, transcript, and source appear in the right place;
  • appointment and payment status are accurate;
  • a task has an owner and due time;
  • current-client messages attach to the correct record;
  • failed writes produce an alert rather than silent loss.

Test category 10: resilience and recovery

Simulate poor audio, long silence, a dropped call, unavailable calendar, integration timeout, failed text delivery, and partial outage. The workflow should degrade safely. It may take a message, offer a callback, create an alert, or explain that a system is temporarily unavailable.


It should not pretend a booking succeeded when it did not. It should not discard information already collected. It should make the failure visible to a named owner.

Run a limited launch

After controlled testing, start with a defined use case such as after-hours calls, overflow, or one practice area. Monitor the first real interactions closely. Do not change several rules at once unless a safety issue requires it, because the team needs to understand which change affected performance.


Review recordings, transcripts, summaries, dispositions, bookings, and downstream records. Use real corrections to improve questions, examples, pronunciation, routing, and escalation.

Measure quality after go-live

Track more than answer rate. The intake metrics that matter include completed intake, qualification, bookings, show rate, follow-up, staff corrections, escalations, and conversion.


For the first weeks, add:


  • boundary violations;
  • incorrect answers;
  • wrong routes;
  • failed transfers;
  • calendar errors;
  • duplicate records;
  • missing fields;
  • integration failures;
  • human takeover rate;
  • defects reopened after a change.


Every defect should have severity, owner, fix, retest scenario, and closure date. A successful launch is a managed operating process, not a one-time demo.

Frequently asked questions

How many calls should a firm test before go-live?

There is no universal number. Cover every major path, rule, language, integration, and failure mode, then repeat critical scenarios after changes. Coverage matters more than a single call count.

Who should participate in testing?

Include a managing attorney, intake owner, receptionist or paralegal, operations or technology owner, and representatives of each practice area or language path.

What should cause a launch delay?

Legal-advice behavior, confidentiality exposure, unreliable emergency routing, incorrect booking or payment confirmation, silent record loss, or repeated failure on a high-volume path should block launch.

Should firms test only business-hours calls?

No. Test after-hours, weekends, overflow, holidays, unanswered transfers, and unavailable calendars.

How should legal-advice requests be tested?

Ask realistic substantive questions and verify that the system states its boundary, avoids an answer, and follows the approved next step.

How do firms test multilingual intake?

Use fluent speakers, natural conversation, accents, spelled names, interruptions, corrections, and language-specific calendars. Do not rely on literal script translation.

What should be inspected in the CRM or practice-management system?

Contact matching, field mapping, summaries, source, appointment, payment, tasks, ownership, duplicates, and visible failure alerts.

Should firms launch every channel at once?

Not necessarily. A limited launch can reduce operational risk, provided the scope and fallback paths are clear.

How often should scenarios be retested?

Retest after material changes to scripts, rules, integrations, calendars, practice areas, staffing, or provider behavior, plus on a recurring quality schedule.

What is the most important go-live metric?

There is no single metric. Response, correct next step, boundary compliance, record quality, and conversion must be reviewed together.


Book a demo with Clerx today.

Share this article:


We use cookies to ensure you get the best experience on our website. For more information, please see our Privacy Policy.