Wednesday, February 4, 2009

Google's approach to the semantic web

Google has never expressed a strong interest in the W3C Semantic Web approach. They are interested in developing systems that can understand the content indexed to a greater degree, however. They have to be, if they take the long view, and Google (still) has the resources to take the long view. The most common reason I have heard from Googlers about why they are not working with RDF and OWL is that they still don't seen enough content out there expressed in these languages. I guess if you are Google, tens or hundreds of billions of triples is still not very much.

IDG news service has a story sketching how Google Researcher Targets Web's Structured Data. This is not directed at data published in mahine understandable form (e.g., in RDF), but on other kinds of structured data accessible on the web.
"Internet search engines have focused largely on crawling text on Web pages, but Google is knee-deep in research about how to analyze and organize structured data, a company scientist said Friday. "There's a lot of structured data out on the Web and we're not doing a good job of presenting it to our users," said Alon Halevy during a talk at the New England Database Day conference at the Massachusetts Institute of Technology,

Halevy was referring in part to so-called "deep Web" sources, such as the databases that sit behind form-driven Web sites like Cars.com or Realtor.com. Google has been submitting queries to various forms for some time, retrieving the resulting Web pages and including them in its search index if the information looks useful.

But the company also wants to analyze the data found in structured tables on many Web sites, Halevy said, offering as an example a table on a Web page that lists the U.S. presidents. And there are reams of those tables -- Google's index turned up 14 billion of them, according to Halevy. He "realized very quickly that over 98 percent of these are not that interesting," but even after significant filtering there remain about 154 million tables worth indexing, he said.
ReadWriteWeb also has a story (Google: "We're Not Doing a Good Job with Structured Data")on that Google is or isn't doing with structured data, including an interesting admission by Google researcher Halevy.
"During a talk at the New England Database Day conference at the Massachusetts Institute of Technology, Google's Alon Halevy admitted that the search giant has "not been doing a good job" presenting the structured data found on the web to its users. By "structured data," Halevy was referring to the databases of the "deep web" - those internet resources that sit behind forms and site-specific search boxes, unable to be indexed through passive means."
For some technical details on the issues and current work, see the paper Google’s DeepWeb Crawl by researchers from Google (including Halevy), UCSD and Cornell published in the Proceedings of VLDB 2009.

Free webinar on the Semantic Web from Dow Jones, Thur 12 Feb 2009

Dow Jones is hosting a free one hour webinar about the Semantic Web, on Thursday 12 February 2009 at 10:00am and again at 2:00pm EST. The webinar, The Semantic Web: Discover, Determine and Deploy, is the first in a tree-part series on the Semantic Web.
"Dow Jones notes that "these days it's critical for organizations to consume, digest, and share news and information. The Semantic Web is no longer ahead of its time and is rapidly changing how organizations keep up with information overload." This webinar is Part I of a series and in it you will learn how Semantic Web Technologies enable you to re-use valuable information to save costs, facilitate easier collaboration and sharing of critical information across your business, and increase search relevancy and surface the most valuable information needed to remain competitive."
The presenters are Christine Connors and Daniela Barbosa, both members of the Dow Jones Enterprise Media Group

The webinar is free but requires registration.

Monday, February 2, 2009

Problems with RDF validator: undecodable data

We're experiencing problems in validating FOAF files. The two identical files served by different servers give different results. The first validates successfully but the other produces an error:
"An attempt to load the RDF from URI 'http://cs.umbc.edu/~ctilmes1/ foaf.rdf' failed. (Undecodable data when reading URI at byte 0 using encoding 'UTF-8'. Please check encoding and encoding declaration of your document.)"
Checking the headers when getting two files shows them to be identical and have reasonable http headers:

% GET http://cs.umbc.edu/~ctilmes1/foaf.rdf | md5sum
74ca1b53dab1591a76517d714686936a -
% GET http://userpages.umbc.edu/~ctilmes1/foaf.rdf | md5sum
74ca1b53dab1591a76517d714686936a -

% HEAD http://cs.umbc.edu/~ctilmes1/foaf.rdf
200 OK
Connection: close
Date: Mon, 02 Feb 2009 15:45:49 GMT
Accept-Ranges: bytes
ETag: "603-dff-461f145761dea"
Server: Apache
Content-Length: 3583
Content-Type: application/rdf+xml
Last-Modified: Mon, 02 Feb 2009 15:33:07 GMT
Client-Date: Mon, 02 Feb 2009 15:45:49 GMT
Client-Peer: 130.85.36.80:80
Client-Response-Num: 1

% HEAD http://userpages.umbc.edu/~ctilmes1/foaf.rdf
200 OK
Connection: close
Date: Mon, 02 Feb 2009 15:45:58 GMT
Accept-Ranges: bytes
ETag: "2a27001e-dff-49871180"
Server: Apache/1.3.33 (Unix) mod_fastcgi/2.4.2 PHP/4.3.10 mod_perl/
1.29 mod_ssl/2.8.22 OpenSSL/0.9.7d
Content-Length: 3583
Content-Type: application/rdf+xml
Last-Modified: Mon, 02 Feb 2009 15:30:08 GMT
Client-Date: Mon, 02 Feb 2009 15:45:58 GMT
Client-Peer: 130.85.24.44:80
Client-Response-Num: 1

We've sent email to the W3C www-rdf-validator list, but if anyon has advice, please let us know.

Jim Hendler on Web 3.0 in Computer, v42n1

Jim Hendler has a short three page article in the January issue of Computer on Web 3.0, aka, to some, anyway, the Semantic Web.
Jim Hendler, Web 3.0 Emerging, Computer, v42n1, pp 88-90, January 2009.
Here's how ACM summarized it in their daily TechNews service:
"Web 3.0 is generally defined as Semantic Web technologies that run or are embedded within large-scale Web applications, writes Jim Hendler, assistant dean for information technology at Rensselaer Polytechnic Institute. He points out that 2008 was a good year for Web 3.0, based on the healthy level of investment in Web 3.0 projects, the focus on Web 3.0 at various conferences and events, and the migration of new technologies from academia to startups. Hendler says the past year has seen a clarification of emerging Web 3.0 applications. "Key enablers are a maturing infrastructure for integrating Web data resources and the increased use of and support for the languages developed in the World Wide Web Consortium (W3C) Semantic Web Activity," he observes.

The application of Web 3.0 technologies, in combination with the Web frameworks that run the Web 2.0 applications, are becoming the benchmark of the Web 3.0 generation, Hendler says. The Resource Description Framework (RDF) serves as the foundation of Web 3.0 applications, which links data from multiple Web sites or databases. Following the data's rendering in RDF, the development of multisite mashups is affected by the use of uniform resource identifiers (URIs) for blending and mapping data from different resources. Relationships between data in different applications or in different parts of the same application can be deduced through the RDF Schema and the Web Ontology Language, facilitating the linkage of different datasets via direct assertions.

Hendler writes that a key dissimilarity between Web 3.0 technologies and artificial intelligence knowledge representation applications resides in the Web naming scheme supplied by URIs combined with the inferencing in Web 3.0 applications, which supports the generation of large graphs that can prop up large-scale Web applications."

Wednesday, January 28, 2009

State of the Semantic Web

Danny Ayers, a well known Semantic Web developer, has the first of a three-part article on the state of the semantic web in the latest issue of IEEE Internet Computing.
Danny Ayers, "Delivered Deliverables: The State of the Semantic Web, Part 1," IEEE Internet Computing, vol. 13, no. 1, pp. 86-89, Jan./Feb. 2009.
He writes in the Nodalities blog:
"Well finally I got around to starting this write-up, and the first instalment has appeared in the excellent IEEE Internet Computing. I foolishly thought I’d be able to cover the main ground in one column, now it seems like I’ll need at least three. In Delivered Deliverables I look mostly at the output of the W3C. The provisional plan is to cover infrastructure & backend tools in part two (with comments of the notion of linked data), and move on to real-world applications in part three. Suggestions are very much welcome."
It's a good summary of the standards that have been developed by the W3C to support the Semantic Web.

Tuesday, January 27, 2009

Semantics-Empowered Social Computing

Here's another interesting looking article from the current issue of IEEE Internet Computing.

Amit Sheth and Meenakshi Nagarajan, Semantics-Empowered Social Computing, IEEE Internet Computing, v13n1, 2 pp 76-80, 2009.
"User-generated textual content on social media has unique characteristics owing to the interpersonal and interactional nature of the communication medium. Web 3.0 applications that aim to automatically create accurate annotations from user-generated content to common reference models will have to invariably deal with the informal nature of this content. In this article, the authors discuss opportunities in addressing challenges posed by this content by supplementing traditional statistical and NLP techniques with domain knowledge."

Semantic Email Addressing

The current issue of IEEE Internet Computing has an article titled Semantic Email Addressing: The Semantic Web Killer App? by Michael Kassoff, Charles Petrie, Lee-Ming Zen, and Michael Genesereth.
"Email addresses, like telephone numbers, are opaque identifiers. They’re often hard to remember, and, worse still, they change from time to time. Semantic email addressing (SEA) lets users send email to a semantically specified group of recipients. It provides all of the functionality of static email mailing lists, but because users can maintain their own profiles, they don’t need to subscribe, unsubscribe, or change email addresses. Because of its targeted nature, SEA could help combat unintentional spam and preserve the privacy of email addresses and even individual identities."
You can get the full pdf here.