Follow-Up to the GenAI Review

Follow-Up to the GenAI Review Case Study: eDiscovery Trends

Last week, I covered a terrific case study on effectively organizing documents with GenAI review. Here’s a follow-up to the GenAI review case study.

Published by EDRM (and available here), the case study was authored by Tara S. Emory, Special Counsel at Covington & Burling LLP, an experienced legal technology lawyer focusing on AI legal applications, eDiscovery, and Information Governance. It involves the practical application and evaluation of GenAI for organizing legal documents in discovery and was conducted by Tara using Relativity’s aiR for Review, outlines a systematic testing framework to assess GenAI’s performance in categorizing over 7,000 conceptually similar documents across nine nuanced sub-issues.

When I read the article, I reached out to Tara with this question as a follow-up to the GenAI review case study:

“I’m interested in understanding more about your quote in the article where you said: ‘Unlike validation-driven TAR protocols often developed with defensibility concerns in mind, our objective was to develop a practical and efficient method to evaluate whether prompts were performing well enough for document organization, prioritization, and issue understanding. Formal validation across nine issues would have been too time-intensive and not aligned with our goals.’ You footnoted the Sedona article you co-wrote with Jeremy Pickens & Wilzette Louis last year (which we covered here).

I’m curious to know more about that decision. Is it because this was a hypothetical situation and not a real litigation case? Would you still apply the principals from the Sedona article if it was a real case or has the thinking changed about how you would approach that situation?”

Tara’s response:

“This was a real case with an urgent need to help our team efficiently find the best evidence, for each of nine sub-issues, which fell under a larger issue that had already been tagged for in prior human review. We wanted to organize the documents so the attorneys could quickly browse for evidence they wanted to cite in their brief.

As noted in the case study, the goal in this case was different from that discussed in the TAR 1 Reference Model article. Here, we were not in a situation requiring defensible recall (finding a reasonable percentage of responsive documents overall), but instead just wanted to efficiently find important documents for each of several issues. We mostly did follow the TAR 1 framework, because it was important to know whether our prompts were working and whether we could improve them. But, we adapted the framework to match the case need, which was finding key evidence efficiently rather than focusing on recall.

The TAR 1 Reference Model is aligned with our process, but the main difference related to the Control Set step. Fn 10 of the Sedona Conference Journal article notes that control sets are not always necessary and may be disproportional to case needs, but they increase likelihood of success because they help you to monitor progress on prompt iterations. We also discussed in that article that the control-set reviewer and prompt writer should be different people so the prompts are not overfit to the control set (written to that set specifically but failing to work on the broader set).

In this case, we skipped the random control set because we did not want to spend valuable reviewer time on random documents. So instead, we worked with a set of key example documents that had already been identified for each issue. We knew this might mean our prompts were more geared to similar key documents instead of the larger set, but that was acceptable for our needs in this matter, and worth the time savings.

Also, to (somewhat) address overfitting, we split the known examples into two test sets. These attorneys had been exposed to documents in both test sets when they first provided us with the examples. which is why it did not completely address overfitting. But, the prompt iterations were done by reviewing results on documents only in Test Set 1. When that was complete, we were able to see how the resulting prompts performed on Test Set 2, which helped us assess overfitting. We also created a small random sample as a very basic validation exercise, and that was Test Set 3.

So this wasn’t a pure implementation of TAR 1 like that needed when production defensibility is the goal, but it worked very well for our specific needs. The iterative testing process was especially useful because it helped us develop our final search strategy, which combined GenAI predictions with targeted search terms and metadata filters.”

Thanks, Tara, for the clarification and additional information! 😊

So, what do you think about this follow-up to the GenAI review case study? Are there questions you have about it? Please share any comments you might have or if you’d like to know more about a particular topic.

Image created using Microsoft Designer, using the term “robot lawyer with nine ‘in’ boxes filled with paper”. Well, sort of. 😉

Disclaimer: The views represented herein are exclusively the views of the author, and do not necessarily represent the views held by my employer, my partners or my clients. eDiscovery Today is made available solely for educational purposes to provide general information about general eDiscovery principles and not to provide specific legal advice applicable to any particular circumstance. eDiscovery Today should not be used as a substitute for competent legal advice from a lawyer you have retained and who has agreed to represent you.


Discover more from eDiscovery Today by Doug Austin

Subscribe to get the latest posts sent to your email.

Leave a Reply