This is a continuation of a blog that I started last week. Suggest you read Part One before this.
Simplified Six Step Review Plan for Small and Medium Sized Cases or Otherwise Where Predictive Coding is Not Used
Here is the workflow for the simplified six-step plan. The first three steps repeat until you have a viable plan where the costs estimate is proportional under Rule 26(b)(1).
Step One: Multimodal Search
The document review begins with Multimodal Search of the ESI. Multimodal means that all modes of search are used to try to find relevant documents. Multimodal search uses a variety of techniques in an evolving, iterated process. It is never limited to a single search technique, such as keyword. All methods are used as deemed appropriate based upon the data to be reviewed and the software tools available. The basic types of search are shown in the search pyramid.
In Step One we use a multimodal approach, but we typically begin with keyword and concept searches. Also, in most projects we will run similarity searches of all kinds to make the review more complete and broaden the reach of the keyword and concept searches. Sometimes we may even use a linear search, expert manual review at the base of the search pyramid. For instance, it might be helpful to see all communications that a key witness had on a certain day. The two-word stand-alone call me email when seen in context can sometimes be invaluable to proving your case.
I do not want to go into too much detail of the types of searches we do in this first step because each vendor’s document review software has different types of searches built it. Still, the basic types of search shown in the pyramid can be found in most software, although AI, active machine learning on top, is still only found in the best.
History of Multimodal Search
Professor Marcia Bates
Multimodal search, wherein a variety of techniques are used in an evolving, iterated process, is new to the legal profession, but not to Information Science. That is the field of scientific study which is, among many other things, concerned with computer search of large volumes of data. Although the e-Discovery Team’s promotion of multimodal search techniques to find evidence only goes back about ten years, Multimodal is a well-established search technique in Information Science. The pioneer professor who first popularized this search method was Marcia J. Bates, and her article, The Design of Browsing and Berrypicking Techniques for the Online Search Interface, 13 Online Info. Rev. 407, 409–11, 414, 418, 421–22 (1989). Professor Bates of UCLA did not use the term multimodal, that is my own small innovation, instead she coined the word “berrypicking” to describe the use of all types of search to find relevant texts. I prefer the term “multimodal” to “berrypicking,” but they are basically the same techniques.
In 2011 Marcia Bates explained in Quora her classic 1989 article and work on berrypicking:
An important thing we learned early on is that successful searching requires what I called “berrypicking.” . . .
Berrypicking involves 1) searching many different places/sources, 2) using different search techniques in different places, and 3) changing your search goal as you go along and learn things along the way. . . .
This may seem fairly obvious when stated this way, but, in fact, many searchers erroneously think they will find everything they want in just one place, and second, many information systems have been designed to permit only one kind of searching, and inhibit the searcher from using the more effective berrypicking technique.
Marcia J. Bates, Online Search and Berrypicking, Quora (Dec. 21, 2011). Professor Bates also introduced the related concept of an evolving search. In 1989 this was a radical idea in information science because it departed from the established orthodox assumption that an information need (relevance) remains the same, unchanged, throughout a search, no matter what the user might learn from the documents in the preliminary retrieved set. The Design of Browsing and Berrypicking Techniques for the Online Search Interface. Professor Bates dismissed this assumption and wrote in her 1989 article:
In real-life searches in manual sources, end users may begin with just one feature of a broader topic, or just one relevant reference, and move through a variety of sources. Each new piece of information they encounter gives them new ideas and directions to follow and, consequently, a new conception of the query. At each stage they are not just modifying the search terms used in order to get a better match for a single query. Rather the query itself (as well as the search terms used) is continually shifting, in part or whole. This type of search is here called an evolving search.
Furthermore, at each stage, with each different conception of the query, the user may identify useful information and references. In other words, the query is satisfied not by a single final retrieved set, but by a series of selections of individual references and bits of information at each stage of the ever-modifying search. A bit-at-a-time retrieval of this sort is here called berrypicking. This term is used by analogy to picking huckleberries or blueberries in the forest. The berries are scattered on the bushes; they do not come in bunches. One must pick them one at a time. One could do berrypicking of information without the search need itself changing (evolving), but in this article the attention is given to searches that combine both of these features.
I independently noticed evolving search as a routine phenomena in legal search and only recently found Professor Bates’ prior descriptions. I have written about this often in the field of legal search (although never previously crediting Professor Bates) under the names “concept drift” or “evolving relevance.” See Eg. Concept Drift and Consistency: Two Keys To Document Review Quality – Part Two (e-Discovery Team, 1/24/16). Also see Voorhees, Variations in Relevance Judgments and the Measurement of Retrieval Effectiveness, 36 Info. Processing & Mgmt 697 (2000) at page 714.
SIDE NOTE: The somewhat related term query drift in information science refers to a different phenomena in machine learning. In query drift the concept of document relevance unintentionally changes from the use of indiscriminate pseudorelevance feedback. Cormack, Buttcher & Clarke, Information Retrieval Implementation and Evaluation of Search Engines (MIT Press 2010) at pg. 277. This can lead to severe negative relevance feedback loops where the AI is trained incorrectly. Not good. If that happens a lot of other bad things can and usually do happen. It must be avoided.
Yes. That means that skilled humans must still play a key role in all aspects of the delivery and production of goods and services, lawyers too.
UCLA Berkeley Professor Bates first wrote about concept shift when using early computer assisted search in the late 1980s. She found that users might execute a query, skim some of the resulting documents, and then learn things which slightly changes their information need. They then refine their query, not only in order to better express their information need, but also because the information need itself has now changed. This was a new concept at the time because under the Classical Model Of Information Retrieval an information need is single and unchanging. Professor Bates illustrated the old Classical Model with the following diagram.
The Classical Model was misguided. All search projects, including the legal search for evidence, are an evolving process where the understanding of the information need progresses, improves, as the information is reviewed. See diagram below for the multimodal berrypicking type approach. Note the importance of human thinking to this approach.
See Cognitive models of information retrieval (Wikipedia). As this Wikipedia article explains:
Bates argues that searches are evolving and occur bit by bit. That is to say, a person constantly changes his or her search terms in response to the results returned from the information retrieval system. Thus, a simple linear model does not capture the nature of information retrieval because the very act of searching causes feedback which causes the user to modify his or her cognitive model of the information being searched for.
Multimodal search assumes that the information need evolves over the course of a document review. It is never just run one search and then review all of the documents found in the search. That linear approach was used in version 1.0 of predictive coding, and is still used by most lawyers today. The dominant model in law today is linear, wherein a negotiated list of keyword is used to run one search. I called this failed method “Go Fish” and a few judges, like Judge Peck, picked up on that name. Losey, R., Adventures in Electronic Discovery (West 2011); Child’s Game of ‘Go Fish’ is a Poor Model for e-Discovery Search; Moore v. Publicis Groupe & MSL Group, 287 F.R.D. 182, 190-91, 2012 WL 607412, at *10 (S.D.N.Y. Feb. 24, 2012) (J. Peck).
The popular, but ineffective Go Fish approach is like the Classical Information Retrieval Model in that only a single list of keywords is used as the query. The keywords are not refined over time as the documents are reviewed. This is a mono-modal process. It is contradicted by our evolving multimodal process, Step One in our Six-Step plan. In the first step we run many, many searches and review some of the results of each search, some of the documents, and then change the searches accordingly.
Step Two: Tests, Sample
Each search run is sampled by quick reviews and its effectiveness evaluated, tested. For instance, did a search of what you expected would be an unusual word turn up far more hits than anticipated? Did the keyword show up in all kinds of documents that had nothing to do with the case? For example, a couple of minutes of review might show that what you thought would be a carefully and rarely used word, Privileged, was in fact part of the standard signature line of one custodian. All his emails had the keyword Privileged on them. The keyword in these circumstances may be a surprise failure, at least as to that one custodian. These kind of unexpected language usages and surprise failures are commonplace, especially with neophyte lawyers.
Sampling here does not mean random sampling, but rather judgmental sampling, just picking a few representative hit documents and reviewing them. Were a fair number of berries found in that new search bush, or not? In our example, assume that your sample review of the documents with “Privileged” showed that the word was only part of one person’s standard signature on every one of their emails. When a new search is run wherein this custodian is excluded, the search results may now test favorably. You may devise other searches that exclude or limit the keyword “Privileged” whenever it is found in a signature.
There are many computer search tools used in a multimodal search method, but the most important tool of all is not algorithmic, but human. The most important search tool is the human ability to think the whole time you are looking for tasty berries. (The all important “T” in Professor Bates’ diagram above.) This means the ability to improvise, to spontaneously respond and react to unexpected circumstances. This mean ad hoc searches that change with time and experience. It is not a linear, set it and forget it, keyword cull-in and read all documents approach. This was true in the early days of automated search with Professor Bates berrypicking work in the late 1980s, and is still true today. Indeed, since the complexity of ESI has expanded a million times since then, our thinking, improvisation and teamwork are now more important than ever.
The goal in Step Two is to identify effective searches. Typically, that means where most of the results are relevant, greater than 50%. Ideally we would like to see roughly 80% relevancy. Alternatively, search hits that are very few in number, and thus inexpensive to review them all, may be accepted. For instance, you may try a search that only has ten documents, which you could review in just a minute. You may just find one relevant, but it could be important. The acceptable range of number of documents to review in Bottom Line Driven Review will always take cost into consideration. That is where Step-Three comes in, Estimation. What will it costs to review the documents found?
Step Three: Estimates
It is not enough to come up with effective searches, which is the goal of Steps One and Two, the costs involved to review all of the documents returned with these searches must also be considered. It may still cost way too much to review the documents when considering the proportionality factors under 26(b)(1) as discussed in Part One of this article. The plan of review must always take the cost of review into consideration.
In Part One we described an estimation method that I like to use to calculate the cost of an ESI review. When the projected cost, the estimate, is proportional in your judgment (and, where appropriate, in the judge’s judgment), then you conclude your iterative process of refining searches. You can then move onto the next Step-Four of preparing your discovery plan and making disclosures of that plan.
Step Four: Plan, Disclosures
Once you have created effective searches that produce an affordable number of documents to review for production, you articulate the Plan and make some disclosures about your plan. The extent of transparency in this step can vary considerably, depending on the circumstances and people involved. Long talkers like me can go on about legal search for many hours, far past the boredom tolerance level of most non-specialists. You might be fascinated by the various searches I ran to come up with the say 12,472 documents for final review, but most opposing counsel do not care beyond making sure that certain pet keywords they may like were used and tested. You should be prepared to reveal that kind of work-product for purposes of dispute avoidance and to build good will. Typically they want you to review more documents, no matter what you say. They usually save their arguments for the bottom line, the costs. They usually argue for greater expense based on the first five criteria of Rule 26(b)(1):
- the importance of the issues at stake in this action;
- the amount in controversy;
- the parties’ relative access to relevant information;
- the parties’ resources;
- the importance of the discovery in resolving the issues; and
- whether the burden or expense of the proposed discovery outweighs its likely benefit.
Still, although early agreement on scope of review is often impossible, as the requesting party always wants you to spend more, you can usually move past this initial disagreement by agreeing to phased discovery. The requesting party can reserve its objections to your plan, but still agree it is adequate for phase one. Usually we find that after that phase one production is completed the requesting party’s demands for more are either eliminated or considerably tempered. It may well now to possible to reach a reasonable final agreement.
Step Five: Final Review
Here is where you start to carry out your discovery plan. In this stage you finish looking at the documents and coding them for Responsiveness (relevant), Irrelevant (not responsive), Privileged (relevant but privileged, and so logged and withheld) and Confidential (all levels, from just notations and legends, to redactions, to withhold and log. A fifth temporary document code is used for communication purposes throughout a project: Undetermined. Issue tagging is usually a waste of time and should be avoided. Instead, you should rely on search to find documents to support various points. There are typically only a dozen or so documents of importance at trial anyway, no matter what the original corpus size.
I highly recommend use of professional document review attorneys to assist you in this step. The so-called “contract lawyers” specialize in electronic document review and do so at a very low cost, typically in the neighborhood of $50 per hour. The best of them, who may often command slightly higher rates, are speed readers with high comprehension. They also know what to look for in different kinds of cases. Some have impressive backgrounds. Of course, good management of these resources is required. They should have their own management and team leaders. Outside attorneys signing Rule 26(g) will also need to supervise them carefully, especially as to relevance intricacies. The day will come when a court will find it unreasonable not to employ these attorneys in a document review. The savings is dramatic and this in turn increases the persuasiveness of your cost burden argument.
Step Six: Production
The last step is transfer of the appropriate information to the requesting party and designated members of your team. Production is typically followed by later delivery of a Log of all documents withheld, even though responsive or relevant. The withheld logged documents are typically: Attorney-Client Communications protected from disclosure under the client’s privilege; or, Attorney Work-Product documents protected from disclosure under the attorney’s privilege. Two different privileges. The attorney’s work-product privilege is frequently waived in some part, although often very small. The client’s communications with its attorneys is, however, an inviolate privilege that is never waived.
Typically you should produce in stages and not wait until project completion. The only exception might be where the requesting party would rather wait and receive one big production instead of a series of small productions. That is very rare. So plan on multiple productions. We suggest the first production be small and serve as a test of the receiving party’s abilities and otherwise get the bugs out of the system.
In this essay I have shown the method I use in document reviews to control costs by use of estimation and multimodal search. I call this a Bottom Line Driven approach. The six step process is designed to help uncover the costs of review as part of the review itself. This kind of experienced based estimate is an ideal way to meet the evidentiary burdens of a proportionality objection under revised Rules 26(b)(1) and 32(b)(2). It provides the hard facts needed to be specific as to what you will review and what you will not and the likely costs involved.
The six-step approach described here uses the costs incurred at the front end of the project to predict the total expense. The costs are controlled by use of best practices, such as contract review lawyers, but primarily by limiting the number of documents reviewed. Although it is somewhat easier to follow this approach using predictive coding and document ranking, it can still be done without that search feature. You can try this approach using any review software. It works well in small or medium sized projects with fairly simple issues. For large complex projects we still recommend using the eight-step predictive coding approach as taught in the TarCourse.com.