Showing posts with label PredictiveCoding. Show all posts
Showing posts with label PredictiveCoding. Show all posts

Friday, May 4, 2012

3 Drawbacks To Predictive Coding

Valora’s Response to LTN article: Take Two: Reactions to 'Da Silva Moore' Predictive Coding Order

What is missing there, and elsewhere, is a discussion of the specific weaknesses of the overall Predictive Coding technique.  Here are just three drawbacks of the technique: 
  1. PC tagging algorithms are not transparent.  No one really knows why the PC engine "chose" the documents it did.  Typically, the “choosing” algorithm is hidden and not disclosed.  All we know is that somehow the document recognized is a lot like another tagged.  
  2. PC has no checks or balances on the skill set, education, consistency or motivations of the seed set coder(s).  The entire Predictive Coding approach assumes that the seed set coder(s) know what they are doing, and that they are correct, consistent and honest.  Would you defend that position, particularly given that the “human being as gold standard" concept has been roundly deflated (see Blair & Maron, Grossman, TREC, etc.)?
  3. Typically, seed set creation and audit sampling for PC use a random sampling technique, the weakest of all types. 
Other sampling techniques (stratified, cluster, panel, etc.) are aware of document attributes and utilize intelligent groupings to create a much stronger, more representative sample for seed set coding and auditing purposes.

Since at present, all Predictive Coding solutions are products, which means they have limited functionality and flexibility for specific case matters, perhaps we should be thinking about the broader picture of Technology-Assisted Review (TAR) as a service – customizable, measurable and transparent.

Wednesday, April 4, 2012

What’s the Difference Between Automated Review and Predictive Coding?

Automated review and predictive coding are often mentioned in the same breath, as synonyms for each other. They are actually different concepts. Predictive coding, in which a topic-expert manually codes a "seed set" of documents (and the software follows suit) is a type of automated review. There are 2 other types.

A second approach to automated review is called Rules-Based Coding, in which a set of rules is created to direct how documents should be coded, very similar to a Coding Manual or a Review Memo that might be prepared for a group of on- or off-shore contract attorneys. The preparation of the Ruleset is typically done by some combination of topic experts, attorneys and technologists. The rules are run on the document population and it is evaluated, tweaked and run again until all parties are satisfied.

The third approach to automated review is called Present & Direct, in which software takes a first, unprompted assessment of the documents and puts forth a graphical representation (pretty charts and diagrams) of what the data contains. This is sometimes called Early Case Assessment or Data Visualization. Once data analysis is presented, the reviewer "informs" the software what he/she wants by batch-tagging key document groupings.

All of these techniques are variations of one another and each has its strengths and weaknesses for use in different types of matters and circumstances. (A topic for a future blog post, clearly!) The point here it is to recognize that Predictive Coding DOES NOT EQUAL Automated Review; it is simply one of several techniques to accomplish it.