Wednesday, April 4, 2012

What’s the Difference Between Automated Review and Predictive Coding?

Automated review and predictive coding are often mentioned in the same breath, as synonyms for each other. They are actually different concepts. Predictive coding, in which a topic-expert manually codes a "seed set" of documents (and the software follows suit) is a type of automated review. There are 2 other types.

A second approach to automated review is called Rules-Based Coding, in which a set of rules is created to direct how documents should be coded, very similar to a Coding Manual or a Review Memo that might be prepared for a group of on- or off-shore contract attorneys. The preparation of the Ruleset is typically done by some combination of topic experts, attorneys and technologists. The rules are run on the document population and it is evaluated, tweaked and run again until all parties are satisfied.

The third approach to automated review is called Present & Direct, in which software takes a first, unprompted assessment of the documents and puts forth a graphical representation (pretty charts and diagrams) of what the data contains. This is sometimes called Early Case Assessment or Data Visualization. Once data analysis is presented, the reviewer "informs" the software what he/she wants by batch-tagging key document groupings.

All of these techniques are variations of one another and each has its strengths and weaknesses for use in different types of matters and circumstances. (A topic for a future blog post, clearly!) The point here it is to recognize that Predictive Coding DOES NOT EQUAL Automated Review; it is simply one of several techniques to accomplish it.

Friday, February 17, 2012

Valora's position on Judge Peck's Commentary in (Da Silva v. Moore) Transcript

The transcript, articles and subsequent brouhaha are overblown. It is obvious that the parties are not particularly debating the efficacy of predictive coding, but rather the proper way to implement it, particularly when there are changes to the population volume or to the relevancy specifications. What is interesting to all of us is this:

Why is there so much wrangling over the size of the seed set and the need to re-seed it with changes?

How should such TAR systems and workflow handle changes? What are the differences between changes to document volume (count) and changes to review specifications?

Why has this ignited the blogosphere, to the point of blatant fact distortion?

There is considerable wrangling over the seed set because of 2 reasons: control and comfort. Judge Peck alludes to the second notion with his distinctions between a statistical sample and a "comfort sample." (His words.) The client service folks in the audience are laughing at this notion as they know full well the mistrust their clients have in "the numbers," while the statisticians present are utterly confounded. Surely, the statistical sample is in fact the most comfortable sample, anything else would be downright uncomfortable!

And, so we get to the second issue here: control. By registering vague, ill-described discomfort with the numbers, the attorneys regain some control by playing on fear. What don't we see? What didn't we get? As Judge Peck points out, it's not about the miniscule "misses," but rather the overall "sure-fires" and the ability to build a case around those.

One of the big problems in this matter is the conflating of two separate types of changes: changes in population size (adding documents) and changes in relevancy scope. These changes are completely independent of one another and should be managed separately. The size change affects the random sampling and so it should be regenerated each time new docs are added. The scope change affects the seed set coding and so it should be regenerated as well to reflect current, up-to-date specifications. In fact, the seed set coding should never be out of step with current specs, or it is obsolete. A far better TAR workflow design than random sampling to generate a seed set and fixed tagging of it is to assume change is part of the litigation and build in mechanisms to manage it. For starters, the seed set should be stratified per the attributes of the then-current document population, rather than random. It should be regenerated every time docs are added or removed. Next, a transparent, easy to edit ruleset should be used to track spec changes and show those over time as documents change their tags. Finally, TAR systems should be priced so that there is no disincentive to make such changes easily.

And finally, why all the hoopla? Because we all know that deep down this is where it is all going. In some ways we are rooting for Recommind (even if we would consider ourselves a competitor) because TAR is the better solution: lower cost, faster and more accurate. Now, if we could be smart about sampling, seed sets and inevitable changes to spec and volume, we'd really have something to shout about.

Friday, October 28, 2011

As I recall...

We at Valora are occasionally called upon to evaluate the results of traditional document review (read: manual, doc-by-doc review efforts), as a sort of post-project audit.  We use the usual precision and recall metrics to determine whether a document should have been tagged at all and whether it was done so correctly.  While both these measures are extremely important, the truth is it’s really all about the recall.  Let me explain.

If you are supposed to mark 1,000 documents as privileged, and your team only tags 200, does their precision on those 200 really matter?  In layman’s terms:  Do you really care how well those 200 docs scored on accuracy when the reviewers missed 800 documents in the first place?

This is not to suggest that there isn’t a crucial role for precision scoring.  There is!  But, not if you blow it on recall first.  In our above example, if the review team indeed found 1,000 documents and marked them privileged, you would surely wish to score the accuracy of those markings.  But, only if the recall fell within 20% +/- of the intended target.  Well below (or above) this mark, precision is useless info.  So, until as an industry we can demonstrate reliable recall on a regular basis, precision will just have to wait its turn.

Wednesday, June 15, 2011

Automated Review Reaches National Economic (& Comedic) Proportions

As a big fan of both The Daily Show and Fareed  Zakaria, imagine my surprise and delight when Jon Stewart’s guest, Zakaria, recently told the live studio audience that Automation was changing the way legal work, specifically discovery, was being conducted today!  Perhaps the most interesting quote was
 “…machines can do things that people used to. There’s now computer programs that can do stuff that lawyers used to be able to do, discovery and things like that.  May not be such a bad thing!”

Well, AutoReveiw has now officially reached national, international really, proportions for both economic analysis and comedic value.  Sounds like we’ve about “crossed the chasm” to me.  For the full June 7, 2011 Stewart-Zakaria interview clip, please see Valora’s website, here.

Wednesday, May 25, 2011

"I Can See Clearly Now..."

Many years ago, early in my career, I worked at software powerhouse, Symantec. As they recently announced their intent to acquire Clearwell, I thought I would pass along my insider’s 2 cents on the topic. I became a "Symantian" in 1996 after they acquired a smaller company with a very hot product in a growing market. At the time the company was called Delrina and they made WinFax. You likely had a copy at some point. It was the innovator in using modems to send faxes and make VOIP calls. (I worked on the telephony end - the cooler part!)

In retrospect, the CW deal sounds very similar to the Delrina one and a good example of how Symantec operates. They bought Delrina at the height of its popularity, paying top dollar. Then they decimated the division, laying off more than half the workforce. (I was not laid off, but asked instead to become GM of a division.) They basically never put another dime into development, QA or marketing and rode the product's margins all the way into the sunset until it was obsolete. I suspect they made a fortune as there were few costs beyond the initial acquisition. They picked up the entire user base and locked out all competition. Within 3 years WinFax as a product was gone and it was completely embedded with other Symantec suite products.

Within 12 months, no one will be charging $300/GB to cull with Clearwell. Every client will have it embedded in their enterprise suite and simply perform the culling themselves. 3rd party vendors beware. The Clearwell party is over and the Relativity party is next.