Showing posts with label Information Governance. Show all posts
Showing posts with label Information Governance. Show all posts

Thursday, June 26, 2014

Why Information Governance is Eclipsing eDiscovery

Everywhere you look right now, Information Governance, “IG,” has taken center stage.  I have had the pleasure of speaking twice in a week on the topic – once to records managers at ARMA’s Northeast Regional Conference and once to litigators at McDermott’s Technology in the Law Symposium.  Why is IG so hot and why is it overtaking the discussion on eDiscovery?

IG is hot, hot, hot
IG is the Next Big Thing because it is a catch-all concept that covers a lot of currently important ground:  compliance, data ethics and breaches, data management and intelligence, workflow, visualization and analytics.  All of these elements play a key part in IG, and always have.  So, why is it hot now?  The biggest factor is the Target data breach of over 70 million customers’ personal data.  The scope of the breach (nearly 25% of all American citizens were affected) and the wall-to-wall media coverage has helped propel responsible data management into the forefront of society’s concerns.  In fact, a Pew Research study in January found that over 50% of Americans are “worried about the amount of personal information available about them…”  The IG train of responsible data management has left the proverbial station and is speeding its way through the legal system, through Wall St., and through consumers’ concerns and buying behavior.  Nothing speaks louder than consumers and their wallets.  For Target, “Satisfactionwith the overall shopping experience was down almost 2 percentage points inMarch, with declines “most acute” among middle-and-upper-income shoppers as late as April, 2014 -four months after the breach was announced.

Why is the IG discussion eclipsing the eDiscovery discussion?
For starters, eDiscovery is old news.  The earliest uses of the phrase stem from 2004, nearly a decade ago, and well before the FRCP changes in late 2006.  Today most litigants, and certainly their outside counsel & advisors are very familiar with its concepts.  In fact, most service providers in legal, lit support or eDiscovery already have a wealth of tools and solutions to choose from.  Need Early Case Assessment?  ESI Processing?  Predictive Coding for Doc Review?  There are a plethora of solutions, all heavily vying for your attention.  The truth is, it’s just not that complicated anymore and the solutions have decreased so much in cost that almost all solutions are accessible to almost all matters.  In short, eDiscovery has become as exciting as word processing or scanning.

But the eDiscovery blahs are only half the reason for the decline in discussion.  The other half is that intelligent IG encompasses eDiscovery.  eDiscovery is subsumed by smart DDC (data, document & content) management, right along with litigation holds, retention policies, workflow routing, exception handling, data breach response and investigations.  Today, eDiscovery is but one of any number of critical activities undergone by major corporations all the time.  It’s just not the fire drill it used to be, and those implementing IG will see to eDiscovery’s needs along the way.  
So, where does that leave us?
Unfortunately, this leaves us woefully and inadequately prepared to handle IG.  The passel of eDiscovery tools do little to solve problems that are much larger than typical litigation matters every imagined.  The the records management side of the house is of little help with their diminished budgets, and dearth of tools available for large-scale data mining and management.  Thus there is promising opportunity for IG-oriented solutions that take the best of both worlds, with an eye towards intelligent DDC management from the outset.  Stay tuned, blog readers, and see where Valora heads next...

Thursday, May 1, 2014

What Do Self-Driving Cars and Documents Have in Common?

In case you missed it, Google had an announcement earlier this week about the rapidly improving reliability of their self-driving cars.  The cars now automatically recognize, pedestrians, trucks, and construction areas and even when a cyclist suddenly veers in front of the car.  Having logged more than 700,000 accident-free miles, it’s an impressive demonstration of a potentially society-altering capability.

So, why mention it here, other than its super-coolness factor?  Because of how it works.  Google’s self-driving cars function because they are taught to recognize patterns.  Patterns of behavior, actions, appearance, movement and trends.  Once a pattern is recognized, the cars’ on-board computer systems run a series of rapid-fire statistical algorithms to determine what is happening (context), and thus what actions the car should take.  Sound familiar?  It should.  This is the same kind of technology behind IBM’s Watson and Valora’s PowerHouse.

Much like Google teaching its cars to recognize a stroller in a crosswalk, Valora teaches PowerHouse to recognize a patent application in Chinese or a break in privilege from an email string.  Google’s vehicles accurately assess and predict traffic behavior patterns almost 100% of the time, much better than human beings.  Valora’s PowerHouse sees similar marks for accuracy and prediction. 

In addition to one day allowing us to text messages or read an e-book while we “drive,” the autonomous vehicles have another enormous advantage:  they better utilize roads, gas and electricity.  These types of benefits have broad-reaching impact beyond whether any one person is using or not using the self-driving car.  The same holds true for autonomous data mining.  Once the document house is in order, everyone benefits from easy, organized search, to intuitive data visualization to automated notifications of significant events.


Too bad I couldn't write this blog entry while on my way to work this morning…

Thursday, March 13, 2014

Interesting Predictions about Data Analytics from Gartner

"Traditional vendors of analytic platforms recognize that in order to expand their reach beyond traditional power users, they must deliver packaged domain expertise and applications to enable self-service by a wider range of users. Service providers are seeking to turn custom project work and domain expertise into repeatable solutions that can be adopted by other organizations more easily.
The result is that end-user organizations selecting analytic applications will have a significantly wider variety of possible providers to evaluate. Organizations evaluating software vendors will almost always find a SaaS version of their packaged applications, and the similarity of product concepts will shift the emphasis of competition to the domain expertise embedded by the vendors into the application. Software vendors will increasingly face a co-opetition situation with their traditional service provider channels, forcing them to augment their own professional service capabilities. Service providers will use packaged applications as an integral part of their customer relationships, implying that there is a greater specialization in the services that they provide."
-Gartner Press Release 12.16.2013

Monday, March 10, 2014

Data Vs. Document Vs. Content

Remember letters?  Typeset documents on official-looking letterhead?  When we communicated primarily via letters, no one wondered what to call the media transmitting information.  It was a Document, plain and simple.  Then came the Internet and websites and eyeballs, and suddenly it was all about Content.  Keeping your content fresh, managing your content, re-using content.  Now it is all about the Data – Big Data, of course.  So, what’s the difference?  Data vs. content vs. document – is there a difference?  In theory, not much, but in practice, yes there is.

Let’s start with Documents.  Documents can be physical or virtual, but they typically have a defined start and end, often delineated by page.  Documents have a specific purpose: they were created by someone, for someone, and they are meant to convey information.  Documents carry with them an air of significance, importance and validity.  That’s why we have phrases like, “Legal documents, financial documents and immigration documents.”  Good examples of documents:  Your tax form, your birth certificate, a receipt from a purchase, your boarding pass.

Content is amorphous.  Though it too can be physical or virtual, it is generally thought of as virtual/electronic in nature only.  Content may or may not have a specific purpose.  It may be written by someone, or sometimes auto-generated.  Content is often not meant to stand on its own, but rather be a supporting player.  Content can be ephemeral, biased and taken out of context.  Because of this, content is not always trusted and carries less validity than documents.  Good examples of content:  blog entries, news, chapters in a book.

Data is virtual.  It is reported, stored or derived from other systems and carries with it a factual and scientific nature.  Data is meant to be bias-free and exist for measurement or tracking purposes.  Good examples of data:  your height and weight, stock prices, bank account balances.

To call information data is to expand on the original intent of what we understand data to be.  However, because our information today is generated and stored electronically, it feels like data, and we (or savvy marketers) have started calling it data.  Thus stored information has becomes data, with all the attached concepts typically assigned to data (factual, bias-free, etc.).  Data, therefore, feels trustworthy and valid – a strong case for managing its exposure.

For more information on the difference between Data and Records, see my article in this month's ARMA newsletter.  When is Data A Record?  (See pages 23-25)

Thursday, October 17, 2013

Why ARMA?

Some people might wonder why a “lit support vendor” would be attending the ARMA National Conference in Las Vegas.  Truth is, Valora’s capabilities have been exceeding “lit support” for a long time.  We find kindred spirits in ARMA, because we are looking at the larger world of corporate documents – for lots of purposes, litigation being just one of them.  In the last 18 months, we have seen tremendous convergence between traditional litigation and eDiscovery with Records Information Management and Information Governance.  In fact, last month I gave a presentation to the NYC chapter of ARMA on “5 Things Litigation Can Teach Records Management and 5 Things You Can Teach Them.”  (Let me know if you’d like a copy of the slides.)  We are going to see more and more of this kind global information management, where litigation is but one use of an organized and controlled data governance strategy.  Watch this space for more on this topic in the weeks to come.

Thursday, October 3, 2013

25 Cool Valora Things

I am often asked, “what’s the coolest thing Valora has ever done?”  That’s a toughie because Valora does a lot of cool things and I would be hard-pressed to pick just one.  Having just gotten yet another totally awesome request yesterday, I decided to compile a list of The 25 Coolest Things Valora Has Ever Done.  

If you think we forgot some, email me at: sserkes@valoratech.com.  And, if you’d like to learn more about any of the real-world scenarios on that list, just email or call.  We’d be happy to share our stories with you (to the extent we are able).

And, finally, here's a bonus cool thing:  an Auto-Generated Word Cloud for the content on this page.

  1. Re-orient and AutoCode documents presented in "mirror writing"
  2. Capture the Japanese "Showa" Date off of documents
  3. Assess long distance spending habits by analyzing multiple years of corporate phone records
  4. Create metadata for (paper) documents from 1901 - 1925, including Near Dupes
  5. Analyze credit card receipts to determine his & hers spending habits for a high profile divorce
  6. Automatically determine if documents are Classified
  7. Identify buildings by address, building number or building name (e.g., "Trump Tower")
  8. Index video files, with generated stills that correspond to key phrases & topics
  9. Uncover an "inappropriate relationship" within standard business communications
  10. Identify likely missing documents from email chains, custodians and shared drives
  11. Code work product documents that included Valora invoices and emails in them (talk about recursive self-reference!)
  12. Determine which applicants were lying on their hiring application
  13. Translate documents to/from Japanese, German, French & English to each of the other 3 languages
  14. Identify bodies of water in documents
  15. Analyze shipping records to identify unusual purchasing behavior
  16. Select "best" versions from multiple reports and coverage of the same event
  17. Audit the results of Onshore Doc Review vs. Offshore Doc Review vs. AutoReview
  18. Determine what type of information was likely underneath document redactions (blackouts)
  19. Identify the cell phone of an NBA player
  20. Match 25,000 index cards with appropriate database records
  21. Redact out ages of minors (no redactions for 21+)
  22. Review documents for 162 unique "Issues"
  23. AutoUnitize a 300,000 page PDF into "logical" documents
  24. Graph potential smuggling routes based on email traffic and news reporting
  25. Index 30 million records in 3 months (that's over 300,000 records every 24 hours)