Wednesday, October 03, 2007

Of Ontologies + Taxonomies

In 2002 -- two years before Tim O'Reilly's famous coining of the term, "Web 2.0," Katherine Adams of the Los Angeles Public Library had already argued that librarians will be an essential piece to the Semantic Web equation. In The Semantic Web: Differentiating Between Taxonomies and Ontologies, Adams makes a few strong arguments that is strikingly ahead of their time. Long before wikis, blogs, and RSS feeds had come to prominence, (5 years ago!) Adams had the foresight to point out the importance of librarians in reply to Berners-Lee et al's vision. Here are Adams main points, all of which I find fascinating based on pre-Web 2.0 knowledge:

(1) Taxonomies: An Important Part of the Semantic Web - The new Web entails adding an extra layer of infrastructure to the current HTML Web - metadata in the form of vocabularies and the relationships that exist between selected terms will make this possible for machines to understand conceptual relationships as humans do.

(2) Defining Ontologies and Taxonomies - Ontologies and taxonomies are used synonymously -- Computer Scientists refer to hierarchies of structured vocabularies as "ontology" while librarians call them "taxonomy."

(3) Standardized Language and Conceptual Relationships - Both taxonomies and ontologies consist of a structured vocabulary that identifies a single key term to represent a concept that could be described using several words.

(4) Different Points of Emphasis - Computer Science is concerned with how software and associated machines interact with ontologies; librarians are concerned with how patrons retrieve information with the aid of taxonomies. However, they're essential different sides of the same coin.

(5) Topic Maps As New Web Infrastructure - Topic maps will ultimately point the way to the next stage of the Web's development. They represent a new international standard (ISO 13250). In fact, even the OCLC is looking to topic maps in its Dublin Core Initiative to organize the Web by subject.

Monday, October 01, 2007

Web 3.0 Librarian

My colleague Dean Giustini and I have collaborated on an article, The Semantic Web as a large, searchable catalogue: a librarian’s perspective. In it, we argue that librarians will play a prominent role in Web 3.0. The current Web is disjointed and disorganized, and searching is much like looking for a needle in the haystack.

It's not unlike the library before Melvil Dewey introduced the idea of organizing and cataloguing books in a classification system. In many ways, we see the parallels here 130 years later. It's not surprising at all to see the OCLC at the forefront in developing Semantic Web technologies. Many of the same techniques of bibliographic control apply to the possibilities of the Semantic Web. It was the computer scientists and computer engineers who had created Web 1.0 and 2.0, but it will ultimately be individuals from library science and information science who will play a prominent role in the evolution of organizing the messiness into a coherent whole for users. Are we saying that Web 2.0 is irrelevant? Of course not. Web 2.0 is an intermediary stage. Folksonomies, social tagging, wikis, blogs, podcasts, mashups, etc -- all of these things are essential basic building blocks to the Semantic Web.

Thursday, September 27, 2007

Libraries and the Semantic Web

Interestingly, not much has been talked about in terms of librarianship and Semantic Web technologies. It's as if there's a gap that can never be bridged: the rustic gatekeeper of books and high-end cutting edge programmer-speak. Quite recently, Jane Greenberg, professor of Library and Information Science at the University of North Carolina at Chapel Hill, has pointed out in Advancing the Semantic Web via Library Functions that there are many similarities between the library and Semantic Web. Here are some:

(1) Each has developed as a response to an abundance of information

(2) Both have mission statements grounded in service, information access, and knowledge discovery

(3) Both have advanced as a result of international and national standards

(4) Both have grown due to a collaborative spirit

(5) Both have become a part of society's fabric (although not so much yet for the Semantic Web)

Monday, September 24, 2007

Four Ways to Look at the Web

The Semantic Web is far from the monolithic artificial intelligent machine which could seemingly process the whim of a user's thoughts. Cade Metz's Web 3.0: Tomorrow's Web, Today offers an excellent and concise glimpse into the different multitude of possibilities of this new Web. Although still in its hyper-conceptual stages, Metz envisions four directions which Web 3.0 could take:

(1) The Semantic Web - A Web where machines can read sites as easily as humans read them. You ask your machine to check your schedule against the schedules of all the dentists and doctors within a 10-mile radius—and it obeys.

(2) The 3D Web - A Web you can walk through. Without leaving your desk, you can go house hunting across town or take a tour of Europe. Or you can walk through a Second Life–style virtual world, surfing for data and interacting with others in 3D.

(3) The Media-Centric Web - A Web where you can find media using other media—not just keywords. You supply, say, a photo of your favorite painting and your search engines turn up hundreds of similar paintings.

(4) The Pervasive Web - A Web that's everywhere. On your PC. On your cell phone. On your clothes and jewelry. Spread throughout your home and office. Even your bedroom windows are online, checking the weather, so they know when to open and close

Tuesday, September 18, 2007

The Seminal on The Semantic

Before Tim O'Reilly, there was Sir Tim Berners-Lee, who often credited as the creator of the Internet. However, what many do not know is that Berners-Lee also preceded many so-called Web 2.0 experts when he had envisioned the Semantic Web (or as many refer to it synonymously as "Web 3.0"). While O'Reilly came along in 2004 to coin Web 2.0, Berners-Lee had long ago created the conceptual foundations in an article co-produced with James Hendler and Ora Lassila, titled The Semantic Web in Scientific American in 2001. Although librarians and information professionals don't need to know the specifics behind the coding technology behind the Semantic Web (that would be asking too much, for much of it is still in development), it is important to have a good grasp of the concepts and a strong understanding of the history and evolution of the Web. Thus, it is important to know that the Semantic Web will be defined by five concepts:

(1) Expressing Meaning - Bring structure to the meaningful content of Web pages, creating an environment where software agents roaming from page to page can readily carry out sophisticated tasks for users. Semantic Web is not a separate Web but an extension of the current one, in which information is given well-defined meaning, better enabling computers and people to work in cooperation.

(2) Knowledge Representation - For Web 3.0 to function, computers must have access to structured collections of information and sets of inference rules that they can use to conduct automated reasoning: this is where XML and RDF comes in, but are they only preliminary languages?

(3) Ontologies - But for a program that wants to compare or combine information across two databases, it has to know what two terms are being used to mean the same thing. This means that the program must have a way to discover common meanings for whatever database it encounters. Hence, an ontology has a taxonomy and a set of inference rules.

(4) Agents - The real power of the Semantic Web will be the programs that actually collect Web content from diverse sources, process the information and exchange the results with other programs. Thus, whereas Web 2.0 is about applications, the Semantic Web will be about services.

(5) Evolution of Knowledge - The Semantic Web is not merely a tool for conducting individual tasks; rather, its ultimate goal is to advance the evolution of human knowledge as a whole. Whereas human endeavour is caught between the eternal struggle of small groups acting independently and the need to mesh with the greater community, the Semantic Web is a process of joining together subcultures when a wider common language is needed.