Thursday, March 14, 2013

Sprechen sie deutsch? Parlez-vous français? You do now!

Remember in Star Trek when the "away team" would encounter a new civilization and be instantly able to communicate with the alien beings by using their handy "UniversalTranslator"? No Tower of Babel in science-fiction! Well, there needn't be one in today's document environment either! With the recent great strides in pattern recognition and content translation, we effectively have a Universal Translator for document and files written in virtually any world language. With support for 65 world languages, Google is to thank for the raw translation effort, while Valora has taken things to the next lelve by implementing the raw capabilities into complex litigation and records management workflows.

For example, on a recent matter we rapidly AutoTranslated documents from 5 foreign languages into English, where they can now be easily understood and managed by the US litigation team. The whole effort took under a week and was 1/10th the cost of manually translating the same material! As with most Automated Solutions Valora offers, a little technology goes a long way! Learn about Valora's AutoTranslation services, by clicking here.

Wednesday, January 9, 2013

12 Tips To Get The Most Out of Technology-Assisted Review ("TAR")

  1. Decide which TAR approach best fits your needs and how you plan to deploy the solution: Do you want the seed set, predictive coding approach or the pattern-matching, rules-based approach? Seed set is good if you don't really know what you want, or you like to "decide on the fly." Rules-based is good if you know what you're looking for and can explain it (similar to how you would train contract attorneys for a large-scale review).

  2. Similarly, do you want TAR as a service or do you want to install a product? Products are typically lower-cost, but less featured or customizable to your specific needs. Services typically cost more, but usually include expert analysis and consulting as part of the package. One consideration in product vs. service is how frequently you encounter a need for TAR and how similar each instance is to the next. Higher frequencies would lead you towards a product, but low similarities across needs would lead you towards services. Remember to include both hard costs (typically dollars outlaid) and soft costs (such as training time and expenses, storage needs, platform support, etc.) in your analysis.

  3. Be realistic about how much you will rely on the coding performed by tool or process and what level of QC you will require. Will you eventually have "eyes on" every document or will you only put "eyes on" subsets of the documents based on relevance or issue criteria? Understanding this early will help you to make the right decisions on pricing, implementation and staffing.

  4. Get comfortable with pricing metrics conversions. Some solutions are sold per document or file, some per GB and some per hour. Here's how to translate between those metrics. Assume: ~ 6,000 docs/files per GB (post processing), and ~ 50 docs/files reviewed per person per hour. Now you can compare pricing for different solutions!

  5. Be explicit about your needs. Do you want a simple yes/no answer for privilege or do you want to know which types of privilege are being invoked? Ex: attorney-client vs. work product. Same for relevance. Is it enough to know simply that a document is relevant or do you need to know why it is relevant (and/or to what degree)?

  6. Map out your workflow and strategy. You (or your client) will need to defend your document production approach. For maximum defensibility, make sure your process is repeatable and transparent. Be wary of TAR solutions that do not disclose why or how propagated decisions are made. Be similarly cautious of solutions that yield different results when different people are "manning" them. Furthermore, make sure that the provider will back you up by providing tangible proof to support the defensibility of the process.

  7. Understand that TAR is an iterative process. The more guidance and feedback you provide, the stronger the results will be. Do not expect the first round to be perfect. You and the systems will both get better over time. As a rough rule of thumb, expect 4-5 iterations.

  8. Think about Exception Handling. Even the best TAR solutions will encounter "problematic" documents from time to time. How will you handle hand-written documents, custom application files or documents written in foreign languages? A good TAR solution should be able to easily identify the docs/files it cannot handle and remove them from the automated processing queue. In other words, don't pay twice for documents that will ultimately need manual processing.

  9. Make good use of Issue Codes. Most sophisticated TAR solutions can handle multiple Issue Codes, providing very helpful tagging and organizational information for Hot or Responsive documents. A consultative TAR solution provider can help you maximize your Issue Codes protocol so that it complements and enhances the production.

  10. Be aware of potential privacy concerns. Many document collections have sensitive or personally identifying information (PII) in their contents that cannot be openly shared. Sophisticated TAR techniques identify, cull and/or automatically redact this information prior to production. TAR approaches can save many hours of manual effort to cleanse data for production.

  11. Choose your solution carefully. Expect that your needs will change over time, both in general and across a single matter. Ideally, the solution provider has full, unfettered access to the TAR engines, so that they can be custom-tailored to your (or your client's) exact circumstances. Be wary of "one size fits all" solutions.

  12. TAR Beyond Document Productions. TAR has uses far beyond review for responsive and privilege. Consider utilizing the techniques when you (or your client) are on the receiving end of a large volume of data. TAR processes can be extremely cost-effective at organizing, cataloging and indentifying trends and data threads in incoming material.

Monday, November 12, 2012

Redaction? There’s an App for That..

Raise your hand if you've either participated in or managed a group of contract attorneys sitting in rows of cubicles (or "stations") redacting documents for production. Whether your tool of choice was a black marker or a cursor, I'm betting it was a thankless, tedious, pain-staking job and you hated it. Wouldn’t it have been nice if you could've taught the computer how to recognize the patterns of PII (private identification information) and have it make the redactions itself?

Well, guess what, Valora heard your anguished cries and we've built an AutoRedaction engine that rivals any group of manual redactors. With blazing speeds, impressive accuracy and astounding savings, our PowerHouseTM system literally autoredacts paper and ESI documents in seconds.

Capitalizing on our extensive experience with pattern-matching technology[1], Valora has built a custom software program that automatically determines the presence of sensitive PII, confidential or privileged information, and then redacts out that information on the image. AutoRedaction takes the form of a black block, with or without a representative stamp, such as "Redacted" or "Employee 123." Redactions can be made permanent, such as for production purposes, or kept temporary, with a technique for "lift and peek," when desired. Redactions can also be made to the underlying text, or on both text and image, if desired.



[1] For more on Probabilistic Hierarchical Context-Free Grammars, see this link on Google Scholar.

Thursday, November 8, 2012

Statistical Pattern Matching Accurately Predicts Presidential Winners and Electoral College Counts, Why Not Privilege and Responsiveness in Litigation?

The technology utilized by political statisticians is finally getting the attention it deserves.  Not because it is partisan, but because it is accurate.  The excellent article in today’s LA Times explains how mathematical models predicted the election outcome well before the first polls had opened. How? By taking the information from numerous sample sets and re-modeling over and over again with different assumptions and weightings. If this sounds a lot like statistical sampling and pattern-matching, then you have been paying attention! The techniques used by the Nate Silvers of the world to classify and label voting patterns are being used right now in litigation to “predict” (or diagnose, if you prefer) for privilege, responsiveness and issues.

At Valora, we call this technique Probabilistic Hierarchical Context-Free Grammars, but others have shortened it to Statistical Pattern Matching, which works just fine. The point is that information about documents (or voter behavior or music choices) has been available for a long time. The only missing piece is the human comfort level with statistics and probabilistic systems.

If the statisticians can call elections, baseball winners and consumer preferences, isn’t it time we let them loose onto document analysis and review? If you’d like a primer on or a demonstration of Probabilistic Hierarchical Context-Free Grammars in litigation, contact us at valoratech.com.

Thursday, August 9, 2012

Valora Technologies CEO, Sandra Serkes, Responds to Craig Ball’s LTN Article on “Next Level” Technology Assisted Review

Original article: Imagining the Evidence

I am pleased to inform both Mr. Ball and the world that the “next level” of TAR, meaning the use of whole documents and populations, rather than selected seed sets, is already here and doing fine.  Rules-Based approaches to TAR are not constrained by the need to create and perfect the selection of a seed set.  Instead, they apply their algorithms and iterations across the entire population, at once, each time.  There is no need for any exemplar document, as the exemplar is the rule itself – thus any document can be evaluated for its “exemplary-ness” and to what degree, where and when.

Furthermore, Mr. Ball discusses the thorny issue of self-interested collection and seed set tagging.  He suggests the opposing party should be the one to set the seed set tags into motion.  This is a step in the right direction.  But, the best approach would be to have both producing and opposing working together to determine relevance – an option easily afforded by a Rules-Based approach.  Rather than having any one party have to sit down and hand-craft a seed set, both sides can agree on the RULES of responsiveness, rather than on whether this document or that one is the better exemplar.  With agreed-upon rules in place, documents are easily assessed not just for yes/no relevance, but also to what degree.

Finally, the notion of “imagining” the documents is very much alive and well in the field of Data Visualization.  We often use this technique in a descriptive way (here’s what your data shows), but it can also very much be used in a proscriptive way (is there anything that looks like this?  How close?).  This concept is very much connected to the current practice of iterating for performance optimization (aka: trading off precision and recall).  TAR systems that utilize the notion of DocType or Attribute templates already have the concept of a “generic” or “iconic” version, essentially an exemplar.  It is trivial to create more templates and use them in a hierarchical manner to test how much a potential document matches the generic exemplars, by relevance priority.