Wednesday, May 14, 2008

Can we beat Google for Web Search?

Sounds like the most difficult question we faced so far? In Today's world we are relied on google so much that we are not in a position to think that we can have a day at work without using google search and also that there is a better search engine than google. Google has created so much hype around and has made us dependent on itself that we benchmark every new search engine against google. Including those who existed before google like yahoo, altavista etc.

But while google is becoming powerful with each new application it releases, and upgradation it does to its search engine there are still few important points which is missing in google search engine which happens to be the core of google.

Most of us who are interested in Search Engine and how it works have read the paper published by Page and Brin on original google search engine architecture and also the initial version of page ranking algorithm. But over last few years they have believed to change the original page rank algorithm. There are few problems with this search engine.
  1. The page rank considers based on Words and the documents.
  2. The google search is based on current web. Whereas the web is growing and evolving with every passing minute. The paradigm of World Wide Web is Persistent Publish and Read. Which holds good to an extent but the web we are looking at today is evolving. We are not in the era of one publisher and many readers but today we have more content producers than readers on web.
  3. The page ranking algorithm uses the index table and the crawler (software) traverses through the links available on page to navigate to next page and so on. The philosophy what google and many other search engines have adopted is to represent the pages as set of nodes (or documents) connected to each other by a static link (HREF). They see it as some sort of tree structure. Whereas the web is not exactly like that. There are pages that do not have any link at all no incoming and outgoing link. Such pages are left behind by google search. An example is my poetry page which is very much hidden from the google search. Though its been on web for almost 3 yrs now. The google crawler managed to reach the main page of my homepage but could not get to the poetry page as there is no link to poetry page from the main page.
  4. We do not maintain a registry which is based on relevance for the web pages outside the page. The google search engine uses the keywords found in the page while indexing. But there are chances that a page which is relavant might not contain the keyword at all.
  5. Though Google is planning to use Latent Semantic Indexing for its next upgrade for page ranking, the accuracy of result is still doubtful.
This was about the problem, but then what is required to beat the google search engine? As discussed in my previous post on similar topic. I stressed the need for a Semantic Search for the web. The semantic search is missing in google search engine and unless that is made available the google search (and for that matter all other search engine) will still give us the irrelavant results (in abundance) when we query them.

Until Next Time... :)

Friday, May 09, 2008

Why Semantic Search?

In my previous post I discussed search engines in general and also how do they build the index table which is the core of any search engine. One thing which became very clear after these studies that the search engines available today are very limited when it comes to functionality. The keyword search does not leave much room for returning the relevant result. In google if we enter Paris Hilton as a search keyword we also get Hilton Hotel in Paris returned as search result that too on the first page. But the search engine there is not able to distinguish that we are not looking Hilton in Paris but Paris Hilton celebrity. On the other hand if we enter Hiton Paris we also get Paris Hilton in our search result. One way or the other the search result we get is not relevant to what we are looking for.

Last night I was reading about Latent Semantic Indexing (LSI) and that did give some hope. I found this page at SEOBook explaining it in a much simpler way about LSI. There are other references as well but this is one page which other than Wikipedia that explains it in a layman's term.

But the million dollar question we are faced with is whether LSI is going to take away the pain of going through irrelevant search results when we query the search engine. In my opinion that is still not very clear. As the algorithm of LSI is still based on the keywords found in the document. And that is not going to take the pain away unless we use the semantic search. But then why semantic search?

Semantic search as most of us know is based on the meanings conveyed by the objects. The term meaning has more depth than it appears from surface. The semantic search is not new, its been there for centuries now. In ancient times philosophers have given the mantra to the world as how to perform the semantic search. Its just that only a handful of people (technologists) today take the pain to read through those literatures. What the current search engines fail today is to restrict the result to what the user wants. We are allowed to input only a bunch of keywords.

In case of Semantic Search the driving factor is context as different terms (or concepts as John F. Sowa describes it) have different meaning or interpretation depending on where they are used. If we build a search engine around these philosophies then we can definitely achieve semantic search (upto a great extent).

Until Next Time... :)

Monday, April 28, 2008

Reusing Ontology

In one of my earlier post I had put emphasis on Why we need a Common Ontology. Over last couple of days while reading through different papers and books I came across few cases which explains why we need a common ontology.

One of the basic idea of Semantic Web is to allow user to reuse an existing ontology if it meets our needs. Alternatively in the worst case we should be able to use a part of it to fulfill our requirements.

At the same time if every user of the web starts to build their own ontology then there would be no common language and shared understanding about anything. There would be no interoperability of any kind between the two agents using those ontologies. There would be no global processing possible either. The exchange of message will not take place and different machines cannot interpret the messages either.

Thus without ontology reuse the very basic idea of Semantic Web is void. Reusing ontology to an extent is even more important than reusing a URI.

Until Next Time.....

Wednesday, April 23, 2008

Building Index Table

In my previous post we discussed about Search Engines in general. We also discussed that there are few basic functionalities of a Search Engine. They are:
  1. Building Index Table
  2. Performing the Search
  3. Building the Result for us
Building Index Table
In this post we will predominantly focus on how the Index Table is being built. The index tables are the heart and brain of a search engine. The process of building index table begins much before the search engine is live. This is an ongoing process which begins much before the search engine is made available and continues till the search engine exists. In a way we can say that the process of indexing determines the quality of result from the search engine.

The indexing is done by a piece of software called Crawler aka Spider. The crawler as the name suggest crawls on the web page and collects virtually all the information it can from the web page. The input to the crawler is the main URL of the web page. Once the crawler receives the URL it performs the following:
  1. Build an Index Table for each and every word on the Page. Since the word may appear more than once in the document it stores the word, the URL and the number of occurences of the word in the document. This is being done for almost all the words found on the page.

  2. Once the crawler is done with building table for each and every word on the page it then navigates to the first link which happens to be a URL again and crawls the new page. ie it performs the similar activity what it did before ie building an index table for each and every word on the page. At this point in time there are two situations possible.

    • It encounters the word that is not part of the current (or previous) document in the index table. So it just adds the new word to the index table along with URL and number of times the word is found in the document.

    • b) The word already existed in the table and in that case it locates the word in the index table and adds the reference to second URL where the word is found to it. Also the number of times the word occurs in the document.

  3. Once the crawler is finished with the current page then it moves on as described in Step 2.

  4. If there are no unvisited link found on the current page then it will go back to the previous page and will start from next link found on the page and will repeat step 2 and 3.
The flaw of this method is that the web is infinite and practically the step 2 - 4 will never finish. The best possible outcome of this procedure is a fraction of web pages are indexed today. Google which is assumed to be the most powerful search engine can index only 1-2% (approx) web pages on the World Wide Web.

This is not the efficient mechanism to build the index table. There has to be a limit where the crawler has to stop going further down the hierarchy and crawl to other pages in the list (in the original page. In the future post of the series I will bring out the discussion on different approaches to crawl the pages. The choice has to be made whether to go for Dephth-First or Breadth-First. For now we assume that the index table is being built and the search engine is ready to perform the search operation.

Performing Search Operation
The index table is used when we type in the keyword to perform search operation. In simplistic term it goes through the index table and searches for the keyword and the document (URL) where it appears and then builds the list of documents to retrieve. But as I said earlier this is the simplest case. The actual result building mechanism has much more to it than just retrieving the documents and presenting it to the user.

In the next post in this series I will bring the perspective of how the result is being built and shown to the user. What affects the page rank and few more details about the same.

Until Next Time... :)

Tuesday, March 25, 2008

Search Engines

After a long gap I am posting something to my blog. Well there has be numerous activities and the most important among all those was getting married last month. The whole February month was filled with travel, meeting family and friends etc. Finally the day when I realized, the holiday was already over and it was time for me to come back to real world. Well I am back now in real world and will be posting something interesting as I discover something on the way of my research on the semantic web.

The most common use of internet today is for searching. The idea is to locate and access information or resources on the web. For example finding out more about the Formula 1 cars etc.

Today the search engines are based on Keywords. They can retrieve documents which contain the given keywords. As long as they given document contains the keyword it will be included in the search result and later shown to the user. The current web then passes on the pain to read and interpret whether the page makes any sense to the user or not. To understand this let us see how search engines are constructed. In this and few more upcoming posts I will be discussing in detail about the search engine and why they function how they function. In this post I will focus primarily on the problem (as described earlier) and some more detail about the search engines.

Today the contains hundreds of millions of web pages. To locate few handful of paged we might be interested in among those hundreds of millions of paged we use search engines like google, yahoo etc. They are the most popular search engines besides others like Altavista, Live Search etc. Inspite of the differences claimed by them a large part of it still remains the same. The fundamental of a search engine building remains the almost the same.

In the future posts will discuss about how the search engines are constructed and why they can do only the keyword search. The next post will be based on creating Index table which is being used by the search engines.

Until Next Time....