Showing posts with label AutoReview. Show all posts
Showing posts with label AutoReview. Show all posts

Thursday, October 3, 2013

25 Cool Valora Things

I am often asked, “what’s the coolest thing Valora has ever done?”  That’s a toughie because Valora does a lot of cool things and I would be hard-pressed to pick just one.  Having just gotten yet another totally awesome request yesterday, I decided to compile a list of The 25 Coolest Things Valora Has Ever Done.  

If you think we forgot some, email me at: sserkes@valoratech.com.  And, if you’d like to learn more about any of the real-world scenarios on that list, just email or call.  We’d be happy to share our stories with you (to the extent we are able).

And, finally, here's a bonus cool thing:  an Auto-Generated Word Cloud for the content on this page.

  1. Re-orient and AutoCode documents presented in "mirror writing"
  2. Capture the Japanese "Showa" Date off of documents
  3. Assess long distance spending habits by analyzing multiple years of corporate phone records
  4. Create metadata for (paper) documents from 1901 - 1925, including Near Dupes
  5. Analyze credit card receipts to determine his & hers spending habits for a high profile divorce
  6. Automatically determine if documents are Classified
  7. Identify buildings by address, building number or building name (e.g., "Trump Tower")
  8. Index video files, with generated stills that correspond to key phrases & topics
  9. Uncover an "inappropriate relationship" within standard business communications
  10. Identify likely missing documents from email chains, custodians and shared drives
  11. Code work product documents that included Valora invoices and emails in them (talk about recursive self-reference!)
  12. Determine which applicants were lying on their hiring application
  13. Translate documents to/from Japanese, German, French & English to each of the other 3 languages
  14. Identify bodies of water in documents
  15. Analyze shipping records to identify unusual purchasing behavior
  16. Select "best" versions from multiple reports and coverage of the same event
  17. Audit the results of Onshore Doc Review vs. Offshore Doc Review vs. AutoReview
  18. Determine what type of information was likely underneath document redactions (blackouts)
  19. Identify the cell phone of an NBA player
  20. Match 25,000 index cards with appropriate database records
  21. Redact out ages of minors (no redactions for 21+)
  22. Review documents for 162 unique "Issues"
  23. AutoUnitize a 300,000 page PDF into "logical" documents
  24. Graph potential smuggling routes based on email traffic and news reporting
  25. Index 30 million records in 3 months (that's over 300,000 records every 24 hours)

Thursday, March 14, 2013

Sprechen sie deutsch? Parlez-vous français? You do now!

Remember in Star Trek when the "away team" would encounter a new civilization and be instantly able to communicate with the alien beings by using their handy "UniversalTranslator"? No Tower of Babel in science-fiction! Well, there needn't be one in today's document environment either! With the recent great strides in pattern recognition and content translation, we effectively have a Universal Translator for document and files written in virtually any world language. With support for 65 world languages, Google is to thank for the raw translation effort, while Valora has taken things to the next lelve by implementing the raw capabilities into complex litigation and records management workflows.

For example, on a recent matter we rapidly AutoTranslated documents from 5 foreign languages into English, where they can now be easily understood and managed by the US litigation team. The whole effort took under a week and was 1/10th the cost of manually translating the same material! As with most Automated Solutions Valora offers, a little technology goes a long way! Learn about Valora's AutoTranslation services, by clicking here.

Wednesday, January 9, 2013

12 Tips To Get The Most Out of Technology-Assisted Review ("TAR")

  1. Decide which TAR approach best fits your needs and how you plan to deploy the solution: Do you want the seed set, predictive coding approach or the pattern-matching, rules-based approach? Seed set is good if you don't really know what you want, or you like to "decide on the fly." Rules-based is good if you know what you're looking for and can explain it (similar to how you would train contract attorneys for a large-scale review).

  2. Similarly, do you want TAR as a service or do you want to install a product? Products are typically lower-cost, but less featured or customizable to your specific needs. Services typically cost more, but usually include expert analysis and consulting as part of the package. One consideration in product vs. service is how frequently you encounter a need for TAR and how similar each instance is to the next. Higher frequencies would lead you towards a product, but low similarities across needs would lead you towards services. Remember to include both hard costs (typically dollars outlaid) and soft costs (such as training time and expenses, storage needs, platform support, etc.) in your analysis.

  3. Be realistic about how much you will rely on the coding performed by tool or process and what level of QC you will require. Will you eventually have "eyes on" every document or will you only put "eyes on" subsets of the documents based on relevance or issue criteria? Understanding this early will help you to make the right decisions on pricing, implementation and staffing.

  4. Get comfortable with pricing metrics conversions. Some solutions are sold per document or file, some per GB and some per hour. Here's how to translate between those metrics. Assume: ~ 6,000 docs/files per GB (post processing), and ~ 50 docs/files reviewed per person per hour. Now you can compare pricing for different solutions!

  5. Be explicit about your needs. Do you want a simple yes/no answer for privilege or do you want to know which types of privilege are being invoked? Ex: attorney-client vs. work product. Same for relevance. Is it enough to know simply that a document is relevant or do you need to know why it is relevant (and/or to what degree)?

  6. Map out your workflow and strategy. You (or your client) will need to defend your document production approach. For maximum defensibility, make sure your process is repeatable and transparent. Be wary of TAR solutions that do not disclose why or how propagated decisions are made. Be similarly cautious of solutions that yield different results when different people are "manning" them. Furthermore, make sure that the provider will back you up by providing tangible proof to support the defensibility of the process.

  7. Understand that TAR is an iterative process. The more guidance and feedback you provide, the stronger the results will be. Do not expect the first round to be perfect. You and the systems will both get better over time. As a rough rule of thumb, expect 4-5 iterations.

  8. Think about Exception Handling. Even the best TAR solutions will encounter "problematic" documents from time to time. How will you handle hand-written documents, custom application files or documents written in foreign languages? A good TAR solution should be able to easily identify the docs/files it cannot handle and remove them from the automated processing queue. In other words, don't pay twice for documents that will ultimately need manual processing.

  9. Make good use of Issue Codes. Most sophisticated TAR solutions can handle multiple Issue Codes, providing very helpful tagging and organizational information for Hot or Responsive documents. A consultative TAR solution provider can help you maximize your Issue Codes protocol so that it complements and enhances the production.

  10. Be aware of potential privacy concerns. Many document collections have sensitive or personally identifying information (PII) in their contents that cannot be openly shared. Sophisticated TAR techniques identify, cull and/or automatically redact this information prior to production. TAR approaches can save many hours of manual effort to cleanse data for production.

  11. Choose your solution carefully. Expect that your needs will change over time, both in general and across a single matter. Ideally, the solution provider has full, unfettered access to the TAR engines, so that they can be custom-tailored to your (or your client's) exact circumstances. Be wary of "one size fits all" solutions.

  12. TAR Beyond Document Productions. TAR has uses far beyond review for responsive and privilege. Consider utilizing the techniques when you (or your client) are on the receiving end of a large volume of data. TAR processes can be extremely cost-effective at organizing, cataloging and indentifying trends and data threads in incoming material.

Thursday, August 9, 2012

Valora Technologies CEO, Sandra Serkes, Responds to Craig Ball’s LTN Article on “Next Level” Technology Assisted Review

Original article: Imagining the Evidence

I am pleased to inform both Mr. Ball and the world that the “next level” of TAR, meaning the use of whole documents and populations, rather than selected seed sets, is already here and doing fine.  Rules-Based approaches to TAR are not constrained by the need to create and perfect the selection of a seed set.  Instead, they apply their algorithms and iterations across the entire population, at once, each time.  There is no need for any exemplar document, as the exemplar is the rule itself – thus any document can be evaluated for its “exemplary-ness” and to what degree, where and when.

Furthermore, Mr. Ball discusses the thorny issue of self-interested collection and seed set tagging.  He suggests the opposing party should be the one to set the seed set tags into motion.  This is a step in the right direction.  But, the best approach would be to have both producing and opposing working together to determine relevance – an option easily afforded by a Rules-Based approach.  Rather than having any one party have to sit down and hand-craft a seed set, both sides can agree on the RULES of responsiveness, rather than on whether this document or that one is the better exemplar.  With agreed-upon rules in place, documents are easily assessed not just for yes/no relevance, but also to what degree.

Finally, the notion of “imagining” the documents is very much alive and well in the field of Data Visualization.  We often use this technique in a descriptive way (here’s what your data shows), but it can also very much be used in a proscriptive way (is there anything that looks like this?  How close?).  This concept is very much connected to the current practice of iterating for performance optimization (aka: trading off precision and recall).  TAR systems that utilize the notion of DocType or Attribute templates already have the concept of a “generic” or “iconic” version, essentially an exemplar.  It is trivial to create more templates and use them in a hierarchical manner to test how much a potential document matches the generic exemplars, by relevance priority.

Friday, May 4, 2012

3 Drawbacks To Predictive Coding

Valora’s Response to LTN article: Take Two: Reactions to 'Da Silva Moore' Predictive Coding Order

What is missing there, and elsewhere, is a discussion of the specific weaknesses of the overall Predictive Coding technique.  Here are just three drawbacks of the technique: 
  1. PC tagging algorithms are not transparent.  No one really knows why the PC engine "chose" the documents it did.  Typically, the “choosing” algorithm is hidden and not disclosed.  All we know is that somehow the document recognized is a lot like another tagged.  
  2. PC has no checks or balances on the skill set, education, consistency or motivations of the seed set coder(s).  The entire Predictive Coding approach assumes that the seed set coder(s) know what they are doing, and that they are correct, consistent and honest.  Would you defend that position, particularly given that the “human being as gold standard" concept has been roundly deflated (see Blair & Maron, Grossman, TREC, etc.)?
  3. Typically, seed set creation and audit sampling for PC use a random sampling technique, the weakest of all types. 
Other sampling techniques (stratified, cluster, panel, etc.) are aware of document attributes and utilize intelligent groupings to create a much stronger, more representative sample for seed set coding and auditing purposes.

Since at present, all Predictive Coding solutions are products, which means they have limited functionality and flexibility for specific case matters, perhaps we should be thinking about the broader picture of Technology-Assisted Review (TAR) as a service – customizable, measurable and transparent.

Wednesday, April 4, 2012

What’s the Difference Between Automated Review and Predictive Coding?

Automated review and predictive coding are often mentioned in the same breath, as synonyms for each other. They are actually different concepts. Predictive coding, in which a topic-expert manually codes a "seed set" of documents (and the software follows suit) is a type of automated review. There are 2 other types.

A second approach to automated review is called Rules-Based Coding, in which a set of rules is created to direct how documents should be coded, very similar to a Coding Manual or a Review Memo that might be prepared for a group of on- or off-shore contract attorneys. The preparation of the Ruleset is typically done by some combination of topic experts, attorneys and technologists. The rules are run on the document population and it is evaluated, tweaked and run again until all parties are satisfied.

The third approach to automated review is called Present & Direct, in which software takes a first, unprompted assessment of the documents and puts forth a graphical representation (pretty charts and diagrams) of what the data contains. This is sometimes called Early Case Assessment or Data Visualization. Once data analysis is presented, the reviewer "informs" the software what he/she wants by batch-tagging key document groupings.

All of these techniques are variations of one another and each has its strengths and weaknesses for use in different types of matters and circumstances. (A topic for a future blog post, clearly!) The point here it is to recognize that Predictive Coding DOES NOT EQUAL Automated Review; it is simply one of several techniques to accomplish it.