Monday, March 10, 2014

Data Vs. Document Vs. Content

Remember letters?  Typeset documents on official-looking letterhead?  When we communicated primarily via letters, no one wondered what to call the media transmitting information.  It was a Document, plain and simple.  Then came the Internet and websites and eyeballs, and suddenly it was all about Content.  Keeping your content fresh, managing your content, re-using content.  Now it is all about the Data – Big Data, of course.  So, what’s the difference?  Data vs. content vs. document – is there a difference?  In theory, not much, but in practice, yes there is.

Let’s start with Documents.  Documents can be physical or virtual, but they typically have a defined start and end, often delineated by page.  Documents have a specific purpose: they were created by someone, for someone, and they are meant to convey information.  Documents carry with them an air of significance, importance and validity.  That’s why we have phrases like, “Legal documents, financial documents and immigration documents.”  Good examples of documents:  Your tax form, your birth certificate, a receipt from a purchase, your boarding pass.

Content is amorphous.  Though it too can be physical or virtual, it is generally thought of as virtual/electronic in nature only.  Content may or may not have a specific purpose.  It may be written by someone, or sometimes auto-generated.  Content is often not meant to stand on its own, but rather be a supporting player.  Content can be ephemeral, biased and taken out of context.  Because of this, content is not always trusted and carries less validity than documents.  Good examples of content:  blog entries, news, chapters in a book.

Data is virtual.  It is reported, stored or derived from other systems and carries with it a factual and scientific nature.  Data is meant to be bias-free and exist for measurement or tracking purposes.  Good examples of data:  your height and weight, stock prices, bank account balances.

To call information data is to expand on the original intent of what we understand data to be.  However, because our information today is generated and stored electronically, it feels like data, and we (or savvy marketers) have started calling it data.  Thus stored information has becomes data, with all the attached concepts typically assigned to data (factual, bias-free, etc.).  Data, therefore, feels trustworthy and valid – a strong case for managing its exposure.

For more information on the difference between Data and Records, see my article in this month's ARMA newsletter.  When is Data A Record?  (See pages 23-25)

Wednesday, November 27, 2013

Data Analytics Steal the Show at DC Technology in the Law Symposium

I was delighted to serve as a panelist at the Technology in the Law Symposium, held earlier this month by the DC Bar Association and McDermott Will & Emery, LLP.  Three panels spoke on the use of predictive technologies and analytics and their use in the courtroom, for eDiscovery and well beyond.  Panelists ranged from outside counsel litigators, to DOJ government attorneys, to service providers and consultants, with Hon. Judge John Facciola presenting the keynote.  The lively, and at times contentious event featured three broad topics:  What the Courts are Saying About Predictive Coding, Predictive Coding Pessimists v. Optimists, and the Use of Data Analytics in Other Areas.

“There’s no greater compliance training than grand jury subpoena.” John Kocoras, Partner, McDermott, Will & Emery


While much of the early discussion around Predictive Coding was, well, somewhat predictable by now, (Key messages:  PC is good, it is proven for doc review, courts are onboard, don’t be reckless with it), the panelists and audience really became animated about the uses of predictive analytics beyond simple relevance review.  Panelists John Kocoras (MWE), Kristian Werling (MWE), Sandra Serkes (Valora) and Kurt Michel (Content Analyst) described scenarios where predictive technologies were being used in multiple corporate settings to assess M&A documents, contracts or financial statements and seek out areas of corporate compliance exposure.  At one point, moderator Jason R. Baron (Drinker Biddle) jokingly asked the panel whether such technologies could accurately predict legal case outcomes.  Answer:  Yes, within reasonable margins of error.

In all, the inaugural Symposium was a roaring success for the 150+ attendees from all over Washington, DC and beyond.  MWE is hoping to expand on their initial success and present several more related symposia in 2014.  Well worth the free attendance!

Thursday, October 17, 2013

Why ARMA?

Some people might wonder why a “lit support vendor” would be attending the ARMA National Conference in Las Vegas.  Truth is, Valora’s capabilities have been exceeding “lit support” for a long time.  We find kindred spirits in ARMA, because we are looking at the larger world of corporate documents – for lots of purposes, litigation being just one of them.  In the last 18 months, we have seen tremendous convergence between traditional litigation and eDiscovery with Records Information Management and Information Governance.  In fact, last month I gave a presentation to the NYC chapter of ARMA on “5 Things Litigation Can Teach Records Management and 5 Things You Can Teach Them.”  (Let me know if you’d like a copy of the slides.)  We are going to see more and more of this kind global information management, where litigation is but one use of an organized and controlled data governance strategy.  Watch this space for more on this topic in the weeks to come.

Thursday, October 3, 2013

25 Cool Valora Things

I am often asked, “what’s the coolest thing Valora has ever done?”  That’s a toughie because Valora does a lot of cool things and I would be hard-pressed to pick just one.  Having just gotten yet another totally awesome request yesterday, I decided to compile a list of The 25 Coolest Things Valora Has Ever Done.  

If you think we forgot some, email me at: sserkes@valoratech.com.  And, if you’d like to learn more about any of the real-world scenarios on that list, just email or call.  We’d be happy to share our stories with you (to the extent we are able).

And, finally, here's a bonus cool thing:  an Auto-Generated Word Cloud for the content on this page.

  1. Re-orient and AutoCode documents presented in "mirror writing"
  2. Capture the Japanese "Showa" Date off of documents
  3. Assess long distance spending habits by analyzing multiple years of corporate phone records
  4. Create metadata for (paper) documents from 1901 - 1925, including Near Dupes
  5. Analyze credit card receipts to determine his & hers spending habits for a high profile divorce
  6. Automatically determine if documents are Classified
  7. Identify buildings by address, building number or building name (e.g., "Trump Tower")
  8. Index video files, with generated stills that correspond to key phrases & topics
  9. Uncover an "inappropriate relationship" within standard business communications
  10. Identify likely missing documents from email chains, custodians and shared drives
  11. Code work product documents that included Valora invoices and emails in them (talk about recursive self-reference!)
  12. Determine which applicants were lying on their hiring application
  13. Translate documents to/from Japanese, German, French & English to each of the other 3 languages
  14. Identify bodies of water in documents
  15. Analyze shipping records to identify unusual purchasing behavior
  16. Select "best" versions from multiple reports and coverage of the same event
  17. Audit the results of Onshore Doc Review vs. Offshore Doc Review vs. AutoReview
  18. Determine what type of information was likely underneath document redactions (blackouts)
  19. Identify the cell phone of an NBA player
  20. Match 25,000 index cards with appropriate database records
  21. Redact out ages of minors (no redactions for 21+)
  22. Review documents for 162 unique "Issues"
  23. AutoUnitize a 300,000 page PDF into "logical" documents
  24. Graph potential smuggling routes based on email traffic and news reporting
  25. Index 30 million records in 3 months (that's over 300,000 records every 24 hours)

Wednesday, July 17, 2013

Specialized Knowledge, Skill, Training and Education

This entry is provided by guest blogger, Aaron Goodisman, Valora’s Chief Technology Officer.

Oh, I feel for D4; I really do. Let me explain:

Law Technology News reports on a case in which D4 Discovery acted as litigation support vendor for both defendant and plaintiff, albeit at different times and performing different functions. Naturally, when defendants Nixon Peabody (working for Kaleida Health) found out, they objected to U.S. Magistrate Judge Leslie Foschio, but he refused to disqualify D4 as a vendor for the plaintiffs.

Sounds like a win for D4, no? As a vendor with many clients in the litigation support space, Valora doesn't like to turn away work any more than the next guy. And, as professionals with over a decade of experience in the legal field, I'm confident that we could maintain appropriate walls of confidentiality between project teams, as D4 asserts they have done.

The problem lies in judge Foschio's rationale for the refusal to disqualify. What the judge essentially said is that D4's scanning and objective coding for Nixon Peabody does not include expertise or consulting, and that it did not expose D4 to any confidential information about the case or Nixon Peabody's case strategy. As an experience scanning and coding provider, this is simply incorrect.

“Objective” coding refers to tagging documents with information that can be objectively determined, without rendering any kind of opinion. In this regard, at least, judge Foschio's rationale makes some sense. That type of information is sufficiently objective that Valora uses software to determine it for most documents. No opinions there.

On the other hand, the design of a scanning and coding project is absolutely a consulting activity: which information is captured for which types of documents, which collections get extra information tagged, which are fast-tracked, which get an extra quality control pass. How the various containment and attachment relationships are captured among documents, folders, binders, boxes. To an experienced litigation support person, those decisions speak volumes about the case.

For proper and accurate scanning and coding to have occurred, D4 had to have access to, and indeed looked at, every single one of the documents in the case, including any that Nixon Peabody later withheld as privileged.

Again, I have no reason to believe that D4 violated their confidentiality responsibilities to either party, nor does it appear that Nixon Peabody is claiming that. Rather, what's happening here is that a judge has said that the services D4 provides do not require “specialized knowledge, skill, training or education.” That's just wrong.

Valora's clients come to us precisely because we provide those things. Kaleida continues to maintain that D4 should have been disqualified from working with the plaintiffs. I'm sure it's standard legal practice, but it feels like somebody's defending the value of such services, at least a little.