Showing posts with label predictive coding. Show all posts
Showing posts with label predictive coding. Show all posts

Wednesday, May 14, 2014

IBM Watson Runs a Food Truck?!

What to do after winning Jeopardy against the world’s best players?  Open a food truck, of course!  Yes, that Watson is now running a food truck, and apparently it creates some truly delicious dishes!  Confused?  Don’t be.  The intelligence behind the Watson engine that successfully answered hundreds of randomized Jeopardy questions is now the creative engine behind a gourmet food truck.  IBM is endeavoring to show that predictive analytics have uses in the most unusual of places!

As with most predictive analytics, there is still an important role for humans to play in providing balance, judgment and expertise.  Watson does  the data-crunching heavy lifting to find interesting and appealing flavor combinations, faster (better?) than human beings can do on their own, and then trained chefs implement the Watson directions.

This hybrid approach should have a familiar ring to it.  Let the software do the hard, data-intensive number-crunching and then marry that output with human skill and finesse.  It’s a winning combination and one that we employ here at Valora every day.  We utilize our analytics, indexing, and rules platform, PowerHouse, to organize, catalog and find relationships in content for us and then we add the human skill, the expertise, to refine the output and do custom things for specific projects.  

Here’s an example:  We run 50,000 emails and attachments through PowerHouse, which quickly finds well over 150 attributes about each item.  Then we ask PH to find important relationships and insights, such as trend data or topic clusters.  From there, we adapt the rules programming to customize the output so it yields middle initials, or zip + 4, or the top 3 issues per document, or whatever it is that any particular customer needs.  Load it up to BlackCat for easy, online review (often by the client’s workforce) and we’re done.  Predictive analytics mastery!


Now, if you’ll excuse me, I think pork belly moussaka sounds amazing!

Wednesday, November 27, 2013

Data Analytics Steal the Show at DC Technology in the Law Symposium

I was delighted to serve as a panelist at the Technology in the Law Symposium, held earlier this month by the DC Bar Association and McDermott Will & Emery, LLP.  Three panels spoke on the use of predictive technologies and analytics and their use in the courtroom, for eDiscovery and well beyond.  Panelists ranged from outside counsel litigators, to DOJ government attorneys, to service providers and consultants, with Hon. Judge John Facciola presenting the keynote.  The lively, and at times contentious event featured three broad topics:  What the Courts are Saying About Predictive Coding, Predictive Coding Pessimists v. Optimists, and the Use of Data Analytics in Other Areas.

“There’s no greater compliance training than grand jury subpoena.” John Kocoras, Partner, McDermott, Will & Emery


While much of the early discussion around Predictive Coding was, well, somewhat predictable by now, (Key messages:  PC is good, it is proven for doc review, courts are onboard, don’t be reckless with it), the panelists and audience really became animated about the uses of predictive analytics beyond simple relevance review.  Panelists John Kocoras (MWE), Kristian Werling (MWE), Sandra Serkes (Valora) and Kurt Michel (Content Analyst) described scenarios where predictive technologies were being used in multiple corporate settings to assess M&A documents, contracts or financial statements and seek out areas of corporate compliance exposure.  At one point, moderator Jason R. Baron (Drinker Biddle) jokingly asked the panel whether such technologies could accurately predict legal case outcomes.  Answer:  Yes, within reasonable margins of error.

In all, the inaugural Symposium was a roaring success for the 150+ attendees from all over Washington, DC and beyond.  MWE is hoping to expand on their initial success and present several more related symposia in 2014.  Well worth the free attendance!

Thursday, May 9, 2013

Technology-Assisted Essay Grading

The NY Times recently reported on the growing use of automated essay grading systems, what we in the legal & records space might call "Technology-Assisted Grading," or "TAG." This is yet another instance of the rest of the world utilizing predictive technologies in conjunction with statistical pattern-matching to create an ultimately subjective judgment of the content of a document. Even more interesting than the fact that MIT & Harvard are making this technology available for free via edX (my alumni donations at work??), is that the higher education community is having the same heated arguments that are occurring right now in the legal arena. Here is the best comment from the piece:

"Although automated grading systems for multiple-choice and true-false tests are now widespread, the use of artificial intelligence technology to grade essay answers has not yet received widespread endorsement by educators and has many critics.."

Sound familiar? It should. This is exactly the argument raging now by outside counsel (playing the part of professors in the article) attempting to hold onto their turf, once considered "untouchable" by technology. While it is true that computers can't "read" either student essays or litigation emails, they can be trained to recognize the salient elements that make the essay strong or the litigation email privileged. Those traits are easily describable as Rules. Either a document fits the Rules, or it doesn't. Nuances are accounted for with confidence scoring and sampling for accuracy (precision & recall). As long as there is sufficient auditing and exception handling, the work quality should be outstanding at a fraction of the time and expense of the purely manual method.

If recent advances have taught us anything, it is that nothing, and certainly no job function, is immutable. It doesn't matter whether the work task is rote (like tightening bolts), cerebral (like computation) or subjective (like analysis), it can all be done by the right algorithms, utilizing the proper training, feedback and statistical sampling.

Furthermore, when subjective work product is automated, society gains impartiality, consistency, speed and reduction of cost for the same services. That allows us to do more, work faster and create better results with fewer resources – the very definition of progress.

It is time to stop fighting the obvious, accept the reality, incorporate the efficiency gain and move on. I'm ready for my essay grade, please.

Wednesday, January 9, 2013

12 Tips To Get The Most Out of Technology-Assisted Review ("TAR")

  1. Decide which TAR approach best fits your needs and how you plan to deploy the solution: Do you want the seed set, predictive coding approach or the pattern-matching, rules-based approach? Seed set is good if you don't really know what you want, or you like to "decide on the fly." Rules-based is good if you know what you're looking for and can explain it (similar to how you would train contract attorneys for a large-scale review).

  2. Similarly, do you want TAR as a service or do you want to install a product? Products are typically lower-cost, but less featured or customizable to your specific needs. Services typically cost more, but usually include expert analysis and consulting as part of the package. One consideration in product vs. service is how frequently you encounter a need for TAR and how similar each instance is to the next. Higher frequencies would lead you towards a product, but low similarities across needs would lead you towards services. Remember to include both hard costs (typically dollars outlaid) and soft costs (such as training time and expenses, storage needs, platform support, etc.) in your analysis.

  3. Be realistic about how much you will rely on the coding performed by tool or process and what level of QC you will require. Will you eventually have "eyes on" every document or will you only put "eyes on" subsets of the documents based on relevance or issue criteria? Understanding this early will help you to make the right decisions on pricing, implementation and staffing.

  4. Get comfortable with pricing metrics conversions. Some solutions are sold per document or file, some per GB and some per hour. Here's how to translate between those metrics. Assume: ~ 6,000 docs/files per GB (post processing), and ~ 50 docs/files reviewed per person per hour. Now you can compare pricing for different solutions!

  5. Be explicit about your needs. Do you want a simple yes/no answer for privilege or do you want to know which types of privilege are being invoked? Ex: attorney-client vs. work product. Same for relevance. Is it enough to know simply that a document is relevant or do you need to know why it is relevant (and/or to what degree)?

  6. Map out your workflow and strategy. You (or your client) will need to defend your document production approach. For maximum defensibility, make sure your process is repeatable and transparent. Be wary of TAR solutions that do not disclose why or how propagated decisions are made. Be similarly cautious of solutions that yield different results when different people are "manning" them. Furthermore, make sure that the provider will back you up by providing tangible proof to support the defensibility of the process.

  7. Understand that TAR is an iterative process. The more guidance and feedback you provide, the stronger the results will be. Do not expect the first round to be perfect. You and the systems will both get better over time. As a rough rule of thumb, expect 4-5 iterations.

  8. Think about Exception Handling. Even the best TAR solutions will encounter "problematic" documents from time to time. How will you handle hand-written documents, custom application files or documents written in foreign languages? A good TAR solution should be able to easily identify the docs/files it cannot handle and remove them from the automated processing queue. In other words, don't pay twice for documents that will ultimately need manual processing.

  9. Make good use of Issue Codes. Most sophisticated TAR solutions can handle multiple Issue Codes, providing very helpful tagging and organizational information for Hot or Responsive documents. A consultative TAR solution provider can help you maximize your Issue Codes protocol so that it complements and enhances the production.

  10. Be aware of potential privacy concerns. Many document collections have sensitive or personally identifying information (PII) in their contents that cannot be openly shared. Sophisticated TAR techniques identify, cull and/or automatically redact this information prior to production. TAR approaches can save many hours of manual effort to cleanse data for production.

  11. Choose your solution carefully. Expect that your needs will change over time, both in general and across a single matter. Ideally, the solution provider has full, unfettered access to the TAR engines, so that they can be custom-tailored to your (or your client's) exact circumstances. Be wary of "one size fits all" solutions.

  12. TAR Beyond Document Productions. TAR has uses far beyond review for responsive and privilege. Consider utilizing the techniques when you (or your client) are on the receiving end of a large volume of data. TAR processes can be extremely cost-effective at organizing, cataloging and indentifying trends and data threads in incoming material.