The Six-Hour Friday Afternoon
A litigation partner at a 22-attorney firm spent every other Friday afternoon reviewing time entries. Six hours, sometimes seven, parsing through narrative descriptions to catch vague entries ("research," "correspondence"), split time between matters, and questionable tenths of an hour.
The work wasn't optional. Clients reject vague billing. State bars audit time records. And associates who don't learn to write clean time entries stay associates longer than they should.
But it was slow. Each entry required judgment: is "reviewed correspondence from opposing counsel" specific enough? Did this associate really spend 2.4 hours on a motion we've filed forty times? The partner kept a running tally of corrections in a spreadsheet, then spent another hour circling back with timekeepers.
What Gets Reviewed and Why It Takes So Long
Most firms review for three things: specificity, allocation, and pattern anomalies.
Specificity means the entry describes what happened in enough detail that a client—or an auditor—can understand the work without asking. "Drafted motion for summary judgment re: plaintiff's statute of limitations argument" passes. "Worked on motion" doesn't.
Allocation catches time that should be split between matters or written off entirely. An associate spends an hour on a brief, realizes halfway through they need to research a new case, and bills the full hour to the brief instead of splitting it.
Pattern anomalies are judgment calls. An associate bills 3.2 hours to review a 12-page contract when the same associate billed 1.8 hours for a similar contract last month. Maybe it was more complex. Maybe they got distracted. The reviewing partner has to decide.
The slowness comes from context-switching. Each entry is a microdecision. Reading the description, remembering the matter, checking if the time makes sense, deciding whether to flag it. Multiply by 400 entries and you've burned a Friday.
Where AI Saves Time—and Where It Doesn't
AI handles the pattern-matching work that doesn't require firm-specific judgment. It flags entries under four words. It catches duplicate time (same task, same matter, billed twice). It identifies descriptions that repeat verbatim across multiple matters, which usually means the associate is copying narrative text without adjusting.
It also flags outliers: time entries that fall two standard deviations outside the norm for similar task codes. A senior associate billing 0.3 hours for a deposition prep gets flagged because the firm's median for that task code is 2.1 hours. The system doesn't know if it's wrong—it just knows it's unusual.
What AI doesn't do: make judgment calls about whether the work was necessary, whether it should be written down, or whether the client will question it. A perfectly specific entry can still be bad billing if the work was redundant or the client didn't authorize it. The reviewing partner still makes those calls.
The value isn't replacing the partner. It's pre-filtering 80% of entries so the partner only looks at the 20% that need human judgment.
The Pilot: Two Weeks, One Matter Type
The firm piloted AI review on commercial litigation matters over a two-week billing cycle. They picked commercial litigation because it had the most entries (roughly 35% of total billed time) and the most variability in associate time patterns.
The AI system reviewed 340 time entries. It flagged 71 for partner review: 43 specificity issues, 19 outliers, 9 possible duplicates. The partner reviewed those 71, made corrections on 38, and approved the rest as acceptable variations.
Total review time: 32 minutes.
The same partner had reviewed a comparable two-week cycle the month before without AI. Review time then: six hours, twelve minutes.
Results: 91% Time Reduction, Zero Billing Errors Found in Audit
The firm ran the pilot for two months, then had their billing auditor review the AI-flagged entries against the full set. The auditor found zero billing errors that the AI had missed. They found two entries the AI had flagged that the partner had approved—both judgment calls, both defensible, but the auditor might have written them down 0.2 hours.
Time saved over eight weeks: 43 hours of partner time. At the firm's internal paralegal rate of $125/hour (the opportunity cost they used for partner admin work), that's $5,375 in recaptured billable capacity. The AI system cost them $600/month. ROI in the first quarter: 223%.
The firm rolled out AI review to all practice groups in month three.
Three Lessons for Firms Considering This
Start with your highest-volume practice group. The AI learns faster when it has more examples, and you'll see ROI faster when you're reviewing hundreds of entries instead of dozens.
Don't skip the human review step. AI flags outliers; it doesn't make decisions. The first month, you'll second-guess every flag. By month two, you'll develop a sense for which flags are worth fifteen seconds and which need two minutes. That calibration only happens if a human is in the loop.
Expect associate resistance. When the AI flags an entry, the associate sees it as criticism. It's not—it's a checklist. But you'll need to explain that at the start, and you'll need to explain it again when the first associate gets a list of twelve flagged entries and assumes they're being dinged for bad billing. They're not. They're being asked to add three words to a time entry.
The partner who ran the pilot now spends Friday afternoons doing partner things. The billing review happens Tuesday mornings and takes less time than the old system took to export the CSV.
Request a demo to see how AI billing review works with your firm's time entry format.