Showing posts with label Document Management. Show all posts
Showing posts with label Document Management. Show all posts

Thursday, May 1, 2014

What Do Self-Driving Cars and Documents Have in Common?

In case you missed it, Google had an announcement earlier this week about the rapidly improving reliability of their self-driving cars.  The cars now automatically recognize, pedestrians, trucks, and construction areas and even when a cyclist suddenly veers in front of the car.  Having logged more than 700,000 accident-free miles, it’s an impressive demonstration of a potentially society-altering capability.

So, why mention it here, other than its super-coolness factor?  Because of how it works.  Google’s self-driving cars function because they are taught to recognize patterns.  Patterns of behavior, actions, appearance, movement and trends.  Once a pattern is recognized, the cars’ on-board computer systems run a series of rapid-fire statistical algorithms to determine what is happening (context), and thus what actions the car should take.  Sound familiar?  It should.  This is the same kind of technology behind IBM’s Watson and Valora’s PowerHouse.

Much like Google teaching its cars to recognize a stroller in a crosswalk, Valora teaches PowerHouse to recognize a patent application in Chinese or a break in privilege from an email string.  Google’s vehicles accurately assess and predict traffic behavior patterns almost 100% of the time, much better than human beings.  Valora’s PowerHouse sees similar marks for accuracy and prediction. 

In addition to one day allowing us to text messages or read an e-book while we “drive,” the autonomous vehicles have another enormous advantage:  they better utilize roads, gas and electricity.  These types of benefits have broad-reaching impact beyond whether any one person is using or not using the self-driving car.  The same holds true for autonomous data mining.  Once the document house is in order, everyone benefits from easy, organized search, to intuitive data visualization to automated notifications of significant events.


Too bad I couldn't write this blog entry while on my way to work this morning…

Thursday, March 13, 2014

Interesting Predictions about Data Analytics from Gartner

"Traditional vendors of analytic platforms recognize that in order to expand their reach beyond traditional power users, they must deliver packaged domain expertise and applications to enable self-service by a wider range of users. Service providers are seeking to turn custom project work and domain expertise into repeatable solutions that can be adopted by other organizations more easily.
The result is that end-user organizations selecting analytic applications will have a significantly wider variety of possible providers to evaluate. Organizations evaluating software vendors will almost always find a SaaS version of their packaged applications, and the similarity of product concepts will shift the emphasis of competition to the domain expertise embedded by the vendors into the application. Software vendors will increasingly face a co-opetition situation with their traditional service provider channels, forcing them to augment their own professional service capabilities. Service providers will use packaged applications as an integral part of their customer relationships, implying that there is a greater specialization in the services that they provide."
-Gartner Press Release 12.16.2013

Monday, March 10, 2014

Data Vs. Document Vs. Content

Remember letters?  Typeset documents on official-looking letterhead?  When we communicated primarily via letters, no one wondered what to call the media transmitting information.  It was a Document, plain and simple.  Then came the Internet and websites and eyeballs, and suddenly it was all about Content.  Keeping your content fresh, managing your content, re-using content.  Now it is all about the Data – Big Data, of course.  So, what’s the difference?  Data vs. content vs. document – is there a difference?  In theory, not much, but in practice, yes there is.

Let’s start with Documents.  Documents can be physical or virtual, but they typically have a defined start and end, often delineated by page.  Documents have a specific purpose: they were created by someone, for someone, and they are meant to convey information.  Documents carry with them an air of significance, importance and validity.  That’s why we have phrases like, “Legal documents, financial documents and immigration documents.”  Good examples of documents:  Your tax form, your birth certificate, a receipt from a purchase, your boarding pass.

Content is amorphous.  Though it too can be physical or virtual, it is generally thought of as virtual/electronic in nature only.  Content may or may not have a specific purpose.  It may be written by someone, or sometimes auto-generated.  Content is often not meant to stand on its own, but rather be a supporting player.  Content can be ephemeral, biased and taken out of context.  Because of this, content is not always trusted and carries less validity than documents.  Good examples of content:  blog entries, news, chapters in a book.

Data is virtual.  It is reported, stored or derived from other systems and carries with it a factual and scientific nature.  Data is meant to be bias-free and exist for measurement or tracking purposes.  Good examples of data:  your height and weight, stock prices, bank account balances.

To call information data is to expand on the original intent of what we understand data to be.  However, because our information today is generated and stored electronically, it feels like data, and we (or savvy marketers) have started calling it data.  Thus stored information has becomes data, with all the attached concepts typically assigned to data (factual, bias-free, etc.).  Data, therefore, feels trustworthy and valid – a strong case for managing its exposure.

For more information on the difference between Data and Records, see my article in this month's ARMA newsletter.  When is Data A Record?  (See pages 23-25)