Showing posts with label AutoCoding. Show all posts
Showing posts with label AutoCoding. Show all posts

Thursday, May 1, 2014

What Do Self-Driving Cars and Documents Have in Common?

In case you missed it, Google had an announcement earlier this week about the rapidly improving reliability of their self-driving cars.  The cars now automatically recognize, pedestrians, trucks, and construction areas and even when a cyclist suddenly veers in front of the car.  Having logged more than 700,000 accident-free miles, it’s an impressive demonstration of a potentially society-altering capability.

So, why mention it here, other than its super-coolness factor?  Because of how it works.  Google’s self-driving cars function because they are taught to recognize patterns.  Patterns of behavior, actions, appearance, movement and trends.  Once a pattern is recognized, the cars’ on-board computer systems run a series of rapid-fire statistical algorithms to determine what is happening (context), and thus what actions the car should take.  Sound familiar?  It should.  This is the same kind of technology behind IBM’s Watson and Valora’s PowerHouse.

Much like Google teaching its cars to recognize a stroller in a crosswalk, Valora teaches PowerHouse to recognize a patent application in Chinese or a break in privilege from an email string.  Google’s vehicles accurately assess and predict traffic behavior patterns almost 100% of the time, much better than human beings.  Valora’s PowerHouse sees similar marks for accuracy and prediction. 

In addition to one day allowing us to text messages or read an e-book while we “drive,” the autonomous vehicles have another enormous advantage:  they better utilize roads, gas and electricity.  These types of benefits have broad-reaching impact beyond whether any one person is using or not using the self-driving car.  The same holds true for autonomous data mining.  Once the document house is in order, everyone benefits from easy, organized search, to intuitive data visualization to automated notifications of significant events.


Too bad I couldn't write this blog entry while on my way to work this morning…

Thursday, October 3, 2013

25 Cool Valora Things

I am often asked, “what’s the coolest thing Valora has ever done?”  That’s a toughie because Valora does a lot of cool things and I would be hard-pressed to pick just one.  Having just gotten yet another totally awesome request yesterday, I decided to compile a list of The 25 Coolest Things Valora Has Ever Done.  

If you think we forgot some, email me at: sserkes@valoratech.com.  And, if you’d like to learn more about any of the real-world scenarios on that list, just email or call.  We’d be happy to share our stories with you (to the extent we are able).

And, finally, here's a bonus cool thing:  an Auto-Generated Word Cloud for the content on this page.

  1. Re-orient and AutoCode documents presented in "mirror writing"
  2. Capture the Japanese "Showa" Date off of documents
  3. Assess long distance spending habits by analyzing multiple years of corporate phone records
  4. Create metadata for (paper) documents from 1901 - 1925, including Near Dupes
  5. Analyze credit card receipts to determine his & hers spending habits for a high profile divorce
  6. Automatically determine if documents are Classified
  7. Identify buildings by address, building number or building name (e.g., "Trump Tower")
  8. Index video files, with generated stills that correspond to key phrases & topics
  9. Uncover an "inappropriate relationship" within standard business communications
  10. Identify likely missing documents from email chains, custodians and shared drives
  11. Code work product documents that included Valora invoices and emails in them (talk about recursive self-reference!)
  12. Determine which applicants were lying on their hiring application
  13. Translate documents to/from Japanese, German, French & English to each of the other 3 languages
  14. Identify bodies of water in documents
  15. Analyze shipping records to identify unusual purchasing behavior
  16. Select "best" versions from multiple reports and coverage of the same event
  17. Audit the results of Onshore Doc Review vs. Offshore Doc Review vs. AutoReview
  18. Determine what type of information was likely underneath document redactions (blackouts)
  19. Identify the cell phone of an NBA player
  20. Match 25,000 index cards with appropriate database records
  21. Redact out ages of minors (no redactions for 21+)
  22. Review documents for 162 unique "Issues"
  23. AutoUnitize a 300,000 page PDF into "logical" documents
  24. Graph potential smuggling routes based on email traffic and news reporting
  25. Index 30 million records in 3 months (that's over 300,000 records every 24 hours)

Wednesday, July 17, 2013

Specialized Knowledge, Skill, Training and Education

This entry is provided by guest blogger, Aaron Goodisman, Valora’s Chief Technology Officer.

Oh, I feel for D4; I really do. Let me explain:

Law Technology News reports on a case in which D4 Discovery acted as litigation support vendor for both defendant and plaintiff, albeit at different times and performing different functions. Naturally, when defendants Nixon Peabody (working for Kaleida Health) found out, they objected to U.S. Magistrate Judge Leslie Foschio, but he refused to disqualify D4 as a vendor for the plaintiffs.

Sounds like a win for D4, no? As a vendor with many clients in the litigation support space, Valora doesn't like to turn away work any more than the next guy. And, as professionals with over a decade of experience in the legal field, I'm confident that we could maintain appropriate walls of confidentiality between project teams, as D4 asserts they have done.

The problem lies in judge Foschio's rationale for the refusal to disqualify. What the judge essentially said is that D4's scanning and objective coding for Nixon Peabody does not include expertise or consulting, and that it did not expose D4 to any confidential information about the case or Nixon Peabody's case strategy. As an experience scanning and coding provider, this is simply incorrect.

“Objective” coding refers to tagging documents with information that can be objectively determined, without rendering any kind of opinion. In this regard, at least, judge Foschio's rationale makes some sense. That type of information is sufficiently objective that Valora uses software to determine it for most documents. No opinions there.

On the other hand, the design of a scanning and coding project is absolutely a consulting activity: which information is captured for which types of documents, which collections get extra information tagged, which are fast-tracked, which get an extra quality control pass. How the various containment and attachment relationships are captured among documents, folders, binders, boxes. To an experienced litigation support person, those decisions speak volumes about the case.

For proper and accurate scanning and coding to have occurred, D4 had to have access to, and indeed looked at, every single one of the documents in the case, including any that Nixon Peabody later withheld as privileged.

Again, I have no reason to believe that D4 violated their confidentiality responsibilities to either party, nor does it appear that Nixon Peabody is claiming that. Rather, what's happening here is that a judge has said that the services D4 provides do not require “specialized knowledge, skill, training or education.” That's just wrong.

Valora's clients come to us precisely because we provide those things. Kaleida continues to maintain that D4 should have been disqualified from working with the plaintiffs. I'm sure it's standard legal practice, but it feels like somebody's defending the value of such services, at least a little.

Thursday, May 9, 2013

Technology-Assisted Essay Grading

The NY Times recently reported on the growing use of automated essay grading systems, what we in the legal & records space might call "Technology-Assisted Grading," or "TAG." This is yet another instance of the rest of the world utilizing predictive technologies in conjunction with statistical pattern-matching to create an ultimately subjective judgment of the content of a document. Even more interesting than the fact that MIT & Harvard are making this technology available for free via edX (my alumni donations at work??), is that the higher education community is having the same heated arguments that are occurring right now in the legal arena. Here is the best comment from the piece:

"Although automated grading systems for multiple-choice and true-false tests are now widespread, the use of artificial intelligence technology to grade essay answers has not yet received widespread endorsement by educators and has many critics.."

Sound familiar? It should. This is exactly the argument raging now by outside counsel (playing the part of professors in the article) attempting to hold onto their turf, once considered "untouchable" by technology. While it is true that computers can't "read" either student essays or litigation emails, they can be trained to recognize the salient elements that make the essay strong or the litigation email privileged. Those traits are easily describable as Rules. Either a document fits the Rules, or it doesn't. Nuances are accounted for with confidence scoring and sampling for accuracy (precision & recall). As long as there is sufficient auditing and exception handling, the work quality should be outstanding at a fraction of the time and expense of the purely manual method.

If recent advances have taught us anything, it is that nothing, and certainly no job function, is immutable. It doesn't matter whether the work task is rote (like tightening bolts), cerebral (like computation) or subjective (like analysis), it can all be done by the right algorithms, utilizing the proper training, feedback and statistical sampling.

Furthermore, when subjective work product is automated, society gains impartiality, consistency, speed and reduction of cost for the same services. That allows us to do more, work faster and create better results with fewer resources – the very definition of progress.

It is time to stop fighting the obvious, accept the reality, incorporate the efficiency gain and move on. I'm ready for my essay grade, please.

Thursday, March 14, 2013

Sprechen sie deutsch? Parlez-vous français? You do now!

Remember in Star Trek when the "away team" would encounter a new civilization and be instantly able to communicate with the alien beings by using their handy "UniversalTranslator"? No Tower of Babel in science-fiction! Well, there needn't be one in today's document environment either! With the recent great strides in pattern recognition and content translation, we effectively have a Universal Translator for document and files written in virtually any world language. With support for 65 world languages, Google is to thank for the raw translation effort, while Valora has taken things to the next lelve by implementing the raw capabilities into complex litigation and records management workflows.

For example, on a recent matter we rapidly AutoTranslated documents from 5 foreign languages into English, where they can now be easily understood and managed by the US litigation team. The whole effort took under a week and was 1/10th the cost of manually translating the same material! As with most Automated Solutions Valora offers, a little technology goes a long way! Learn about Valora's AutoTranslation services, by clicking here.