Sunday, March 15, 2009

Ontology Summit 2009: Toward Ontology-based Standards

A two day event, Ontology Summit 2009: Toward Ontology-based Standards, will be held 6-7 April 2009 at NIST in Gaithersburg MD. The Summit is co-organized by NIST and a number of other organizations.
"This summit will address the intersection of two active communities, namely the technical standards world, and the community of ontology and semantic technologies. This intersection is long overdue because each has much to offer the other. Ontologies represent the best efforts of the technical community to unambiguously capture the definitions and interrelationships of concepts in a variety of domains. Standards -- specifically information standards -- are intended to provide unambiguous specifications of information, for the purpose of error-free access and exchange. If the standards community is indeed serious about specifying such information unambiguously to the best of its ability, then the use of ontologies as the vehicle for such specifications is the logical choice. Conversely, the standards world can provide a large market for the industrial use of ontologies, since ontologies are explicitly focused on the precise representation of information. This will be a boost to worldwide recognition of the utility and power of ontological models. The goal of this Ontology Summit 2009 is to articulate the power of synergizing these two communities in the form of a communique in which a number of concrete challenges can be laid out. These challenges could serve as a roadmap that will galvanize both communities and bring this promising technical area to the attention of others."
The meeting is free, but advanced registration is required. You can also register to participate remotely.

Saturday, March 14, 2009

Video from Tim Berners-Lee 2009 TED talk on linked data

Here is the video of the talk that Tim Berners-Lee gave at the TED2009 conference on linked data.
You can see the slides that TBL used on the W3C site.

I may have missed it, but I don't think he mentioned the phrase "Semantic Web" once during the 16 minute talk.

Monday, March 2, 2009

Ian Davis code{4}lib keynote: data outlasts code

Ian Davis, CTO of Talis, posted the slides from his code4lib2009 keynote talk on slideshare. If you love something... set it free gives a very nicely done description of the motivation behind and hopes for the Semantic Web.

Code{4}lib is a conference series and community focused on the intersection of libraries, technology, and the future. code4lib2009 was held this week in Providence, hosted by the Brown University Library.

Ian's talk contained three conjectures, the first of which I especially liked:
  • Conjecture 1: Data outlasts code
  • Conjecture 2: There is more structured data in the world than unstructured
  • Conjecture 3: Most of the value in our data will be unexpected and unintended
(h/t Danny Ayers)


Tuesday, February 17, 2009

Tim Berners-Lee's map of the Web world

Back in 2007, Tim Berners-Lee made this amusing map. It might be a useful reference as you make your way from the World Wide Web toward the Sea of Interoperability while trying to avoid some of the dark and dangerous places where giants and dragons are said to dwell.

Thursday, February 12, 2009

Yahoo BOSS exposes structured data in RDF

This could be a big step toward the "web of data" vision of the Semantic Web.

Yahoo announced (Accessing Structured Data using BOSS that their BOSS (Build your Own Search System) will now support structured data, including RDF.
"Yahoo! Search BOSS provides access to structured data acquired through SearchMonkey. Currently, we are only exposing data that has been semantically marked up and subsequently acquired by the Yahoo! Web Crawler. In the near future, we will also expose structured data shared with us in SearchMonkey data feeds. In both cases, we will respect site owner requests to opt-out of structured data sharing through BOSS."
Yahoo\'s BOSS to support RDF data
Here's how it works:
  • Sites use microformats or RDF (encoded using RDFa or eRDF) to add structured data to their pages
  • Yahoo's web crawler encounters embedded markup and indexes the structured data along with the unstructured text
  • A BOSS developer specifies "view=searchmonkey_rdf" or "view=searchmonkey_feed" in API requests
  • BOSS's response returns the structured data via either XML or JSON
Yahoo's SearchMonkey only acquires structured data using certain microformats or RDF vocabularies. The microformats supported are hAtom, hCalendar, hCard, hReview, XFN, Geo, rel-tag and adr. RDF vocabularies handled include Dublin Core, FOAF, SIOC, and "other supported vocabularies". See the appendix on vocabularies in Yahoo's SearchMonkey Guide for a full list and more information.

A post on the Yahoo search blog talks about this and other changes to the BOSS service and includes a nice example of the use of structured data encoded using microformats from President Obama’s LinkedIn page.

microformatted data on President Obama\'s linked in page

Sunday, February 8, 2009

Tim Berners-Lee talks on linked data at TED 2009

Tim Berners-Lee gave a talk at the TED2009 conference on linked data -- one of the newest and most interesting ideas to emerge from efforts to realize the Semantic Web vision.

Here's a summary of Sir Beerners-Lee's from a post by Gigaom, Highlights from TED: Tim Berners-Lee, Pattie Maes, Jacek Utko. I'm looking forward to being able to see his talk online soon.
"Founder of the web Tim Berners-Lee spoke of the next grassroots communication movement he wants to start: linked data. Much in the way his development of the web stemmed out of the frustrations of brilliant people working in silos, he is frustrated that the data of the world is shut apart in offline databases.

Berners-Lee wants raw data to come online so that it can be related to each other and applied together for multidisciplinary purposes, like combining genomics data and protein data to try to cure Alzheimer’s. He urged “raw data now,” and an end to “hugging your data” — i.e. keeping it private — until you can make a beautiful web site for it.

Berners-Lee said his dream is already on its way to becoming a reality, but that it will require a format for tagging data and understanding relationships between different pieces of it in order for a search to turn up something meaningful. Some current efforts are dbpedia, a project aimed at extracting structured information from Wikipedia, and OpenStreetMap, an editable map of the world. He really wants President Obama, who has promised to conduct government transparently online, to post linked data online."
You can see the slides that TBL used on the W3C site.

Big data, linked or not

The Data Evolution blog has an interesting post that asks Is Big Data at a tipping point?. It's suggests that we may be approaching a tipping point in which large amounts of online data will be interlinked and connected to suddenly produce a whole much larger than the parts.
"For the past several decades, an increasing number of business processes– from sales, customer service, shipping - have come online, along with the data they throw off. As these individual databases are linked, via common formats or labels, a tipping point is reached: suddenly, every part of the company organism is connected to the data center. And every action — sales lead, mouse click, and shipping update — is stored. The result: organizations are overwhelmed by what feels like a tsunami of data. The same trend is occurring in the larger universe of data that these organizations inhabit. Big Data unleashed by the “Industrial Revolution of Data”, whether from public agencies, non-profit institutes, or forward-thinking private firms."
I expected that the post would soon segue into a discussion of the Semantic Web and maybe even the increasingly popular linked data movement, but it did not. Even so, it sets up plenty of nails for which we have a an excellent hammer in hand. I really like this iceberg analogy, by the way.
"At present, much of the world’s Big Data is iceberg-like: frozen and mostly underwater. It’s frozen because format and meta-data standards make it hard to flow from one place to another: comparing the SEC’s financial data with that of Europe’s requires common formats and labels (ahem, XBRL) that don’t yet exist. Data is “underwater” when, whether reasons of competitiveness, privacy, or sheer incompetence it’s not shared: US medical records may contain a wealth of data, but much of it is on paper and offline (not so in Europe, enabling studies with huge cohorts)."
The post also points out some sources of online data and analysis tools, some familiar and some new to me (or maybe just forgotten.)
"Yet there’s a slow thaw underway as evidenced by a number of initiatives: Aaron Swartz’s theinfo.org, Flip Kromer’s infochimps, Carl Malamud’s bulk.resource.org, as well as Numbrary, Swivel, Freebase, and Amazon’s public data sets. These are all ambitious projects, but the challenge of weaving these data sets together is still greater."