Showing posts with label google. Show all posts
Showing posts with label google. Show all posts

Tuesday, September 09, 2008

It's Not the Content Silly, but the Metadata You Should be Listening to

I have been biting my bottom lip all day- at first because i really did not feel like commenting on it and specifically because i didn't know the 'facts'. But after noticing a swollen lip, i decided to go ahead.

Since i happen to work for a media provider (Dow Jones) whose model is certainly currently focused on providing 'authoritative' premium content- something that many clients are willing to pay a lot of money for access to (so things like today are not a common occurrence)- i knew that most likely any post would be blatantly pro-premium subscriptions- because trust me if someone had done a check on a Dow Jones news feed product they would have seen that the story was 'bogus' with little effort. So a story like this perhaps makes us drool because it provides a perfect case story on why 'free' is not better- certainly not when millions of dollars are at stake.

From a Wired Blog a good explanation of what happened:
the article in the Sun Sentinel's archive had no date on it. But when Google's spider grabbed it, it assigned a current date to the piece, which then resulted in the article being placed in the top results of Google News. When the employee from Income Securities Advisors ran a Google search on "2008 bankruptcies," the old United Airlines story appeared as the top link in the results, with a September 6, 2008 date on it. (Google has now released a screenshot that shows the UAL story as it appeared on the Sun Sentinel web site. The only date in the screenshot is September 7, 2008, the date Google accessed the page. There is no date under the story's headline to indicate when it was published.) At 11 am Monday, the employee added the story to a feed that is included in a Bloomberg subscription service and within minutes, 15 million shares of United Airlines stock had been sold before trading on the stock was halted.

But although this of course is a compelling story to only trust authoritative sources and premium content aggregators, this story is not only about 'free' content because the traders (human or machines i may ask?) acted on a news story from a reputable and costly service- Bloomberg.

So who is to blame? Well of course many are talking about it in both main stream media and blogs and sure things like this have happened before but I think that the simplest answer is to ask why the person who pressed the send button to Bloomberg didn't vet out the story. Perhaps it was early in the morning (ok 11am is not that early but let's go with that excuse....) and he got in late the night before thanks to a delayed United flight and all the recent news about the airline industry losing money and charging $15 for a pillow was enough to believe the story and pass it along as a true story without any vetting.

But the interesting part to me is the metadata associated with the news story- because essentially that was the technical culprit- the article did not have metadata (in this case publication date) to tell Google that it was an article from 2002 that had been republished on their website. The problem is that there really is no standard to provide that information that online news providers adhere to and as more of their archives that were traditionally only available through premium aggregators that normalize the content, come on-line for 'free' more unknowns start to be thrown at online news services like Google News.

One of the core benefits of aggregator premium services (e.g. Factiva from Dow Jones, LexisNexis) is the normalization of the content from 'trusted' sources. This ensures that publication data- sometimes down to the millisecond is provided and the consuming application whether it is a trading system, a news portal or an alerting sms message sent to the banker on the run- gets it right.


Image|Flickr|arimoore

Thursday, April 03, 2008

Semantic Web Experimental Mashup

Via Peter Reiser's blog a pointer to this very cool experimental mashup using Wikipedia, Open Calais, Goggle and Amazon that demonstrates how to use semantic based term extraction and the Amazon API to search for relevant books for a specific topic like for example i used Dow Jones which brought back books on the Wall Street Journal which is one of the Dow Jones properties.


From Peter's Blog: How it works

1. The search input is sent to wikipedia.org
2. The respective wikipedia page is sent to the Open Calais service to extract the terms
3. The extracted terms are sent to google, and get enriched by related terms using the google labs service "google suggest".
4. the terms are sent to the Amazon API and the relevant books from Amazon are displayed

A simple use case for this could be a blog widget that would semantically extracted information from posts, like for example on this post an Amazon widget would present books on the Semantic Web, Google and APIs with my affiliate ID embedded so if a reader wanted to purchase books related to the subject of the specific post they could.

Monday, January 28, 2008

Are we tired of Simple and Advanced Searching features?

Are we getting bored with the little simple search box?

Today's buzz about Google's new search views reminded me to go and read this article from Boxes and Arrows that recently came across my radar about Advancing Advanced Search. If you are interested in the various ways 'advance' search is approached the article itself is quite useful as are many of the comments that follow. There really is no 'one way' to present advanced searching capabilities to users as it really depends on the users, the data being searched and the technologies but go take a look at some of the examples Steven Turbek highlights.


Over at the Google Labs Experimental page they are always highlighting different things that they are working on and i check in often- but honestly i still stick with the simple google search box (but yes i have learned a lot of advanced commands that i use in the search box because i am just geeky like that). However, every once in a while -quite randomly it seems- i get a glimpse into some new features that they are trying out.

Today's announcement of some new views had an interesting Google Map View search feature which could be useful for searching people, companies, events and places.

From the Google blog post:
Map view Suppose you're scouring the web trying to find out about biology conferences happening in your state. Or you'd like to sit back and enjoy some jazz around town. This information is on the web and accessible through regular web search, but probably spread out over many sites and pages. Unless one of these pages has a map, it might be hard to visualize all the locations at once. Map view solves this problem by plotting some of the key locations contained in your web results onto a map.

Here is a sample search for Web 2.0 conferences in the Bay Area. Ok results but could be better- but it is not all Google's fault.

OK great so what would really make this powerful?

You know how i have been mentioning Microformats and RDFa and other types of data standards that are about adding semantic information to web pages without affecting what the page looks like to the naked eye?

Well as the practice of embedding semantic information is adopted throughout the web community what impact can it have? Well for starters it will make a search tool like what Google Map View is trying to do be a quite powerful way to search for things like conferences, events etc....So for that, the Google is on the map with this one.

Friday, January 18, 2008

Report on the Information Behavior of the Researcher of the Future

I know you are all scouring the web for good reading material for the weekend and your first stop was here (ha!). So this will probably be good weekend reading if you are interested in the topic of Search as it relates to research. i just browsed through it and saved it to my 'to read weekend' list so i will update this post with my thoughts after reading in-depth.

Via Techmeme, a pointer to the Arstechnica blog which posts about a new briefing report on the habits of the 'Google' generation titled "Information Behavior of the Researcher of the Future".

Information Professionals, have been discussing this topic for a while especially in the corporate research space when users are given access to specialized databases. Although Google has done a lot to make everyone a 'searcher'- i am not convinced that it makes people good searchers and especially not good researchers. But that is really not Google's 'fault'. Google can be a sophisticated search engine if for example the advanced features are used- users are just not trained to take advantage of them. They can certainly try to educate users- for example working with Libraries to provide teaching resources where early search education takes place.

I would read it right now but lunch time is over and i have some work to do before this evening so i can enjoy myself at the Crunchies which by the way you may also be able to view via video stream this evening.

Image: from Arstechnica post but is actually the cover page of the report.

Wednesday, November 28, 2007

But is the Google Intranet useful?

Being a fan of intranets (how geeky does that sound?) i of course had to click through on on this "What the Google Intranet Looks like" article via Techmeme. A good "investigative" piece more so then an in depth look at how an intranet is used in an enterprise, but i guess it is always hard not to sensationalize the internal workings of Google.

I wonder however if 1 in 3 Google employees like the results of this study find the intranet “not useful” ?

If you want some decent case studies on Enterprise intranets IntranetBlog.com provides a good resource that gives you a peak into the intranets of multinational global companies. Many intranet managers are also on the 'talk' circuit sharing their pains and successes and i use tools like SlideShare to find examples of intranet case studies so i can see what companies in different sectors are up to.

Friday, September 21, 2007

Google Reader new features and potential

One of the gems that i came home to was the release of Google Reader's search feature which many of us had been begging for. The ability to search through all your feeds (from the moment you subscribed to them) is awesome and much needed- i am always looking for that blog post "i read the other day". Another good use that i have already put into practice (principally because i have thousands of unread feeds at this point) is to do quick searches to see who is writing about topics that i might be interested in currently.

I would like to see other features beyond more advance search features, like the ability to save and share searches and APML support.

Steve Rubel posted an interesting post on How to Data Mine Google Reader Feeds for Trends, which is yet another way that these new features can be used.
Rubel via Jeremiah's Twit on Rubel's post.

Wednesday, August 08, 2007

Letters to the Editor- but wait the Editor didn't write the story

Lot's of chatter today around Google's new feature that will allow only 'the people in the news article' to add comments to articles that are aggregated in Google News. Here are two examples that the Google Blogoscoped blog pointed out.

When i saw it this morning i thought it was an interesting use of social media and still believe so-but as i think about it and read a bit further about exactly how it will work there certainly seems to be some potential issues. Things like what Gabe Rivera left as a comment on TechCrunch that Google doesn't allow other aggregators to crawl their site although now they are now hosting original content and probably should. Since Google gets its content by aggregating across thousands of online news sites and there is constant battles about the right to do that- they should probably open up this orginal content back- but who knows maybe the content will be useless in the long run?

i think that it is interesting that Google is going to try to be some sort of 'editor' of the global online news. Thinking about traditional media there are usually two types of ways that someone 'associated' with the article can respond and let the audience hear their voice. The first is to contact the media outlet and request a correction or clarification. I tend to browse these section of the papers and there always seems to be something that needs correction sometimes important information. The second is to write a letter to the editor. Both of these however become physically separated from the original source- either printed in the next day publication (for newspapers but could be months for other types of publications) and essentially the editor of that publication who is ultimately responsible for what they printed has the last word on what they choose to publish on behalf of the person responding.

So will it be scalable - and will Google have to open it up to let the original sources get to that content? Either way- kudos to Google for pushing the envelope and looking a user generated content differently then what other news sites are doing.

Tuesday, June 12, 2007

Google PowerPoint Viewer released for Gmail

For the last few weeks people have been talking about Google PowerPoint Viewer in Gmail and this morning it seems to have gone live. Many times when i am on the road, i send my slide decks to my Gmail account for easy access at the client site (i used to use my memory stick for that purpose). Sometimes, for example if you are on a loaner client laptop or in the training room (or your client doesn't use Microsoft products!) it can be a hassle to have to install software or transfer the PPT to another format. Well thanks to Google this is no longer an issue.

I will definitely use the Slideshow view to take a client through a slide deck but after testing it this morning i still think they need to enhance it, for example transitions within slides are lost and sometimes in the slide decks i use (well the exciting ones!) that is essential. I would also like to see some of the functionality i posted about with Google's acquisition of GapMinder into their Slideshow/presentation tools in order to create dynamic and compelling slidedecks quickly.

Wednesday, May 30, 2007

Google Reader Offline- thanks about time someone did it

It is so nice to know that the folks over at Google are reading my blog (well actually i don't know that i just think everyone one of their releases is about me). With today's announcement of Google Gears that will enable developers to create offline web applications using JavaScript APIs. With the Google Reader offline mode is the first release they have managed to solve yet another one of my problems.

Last December as i was packing for a trip to Europe, i posted about the fact that i had a process in place to get my podcasts and videos all set to be offline but i had no way to go offline with the blogs i read. My short term fix until today was that as i get on a plane i open multiple tabs on my browser and then look through them offline. I even had someone tap me on the shoulder on a plane to Florida a couple months back asking me if i had Internet access on the plane- i didn't- i had just opened multiple tabs to read through when i went offline.

Well i used to have to make do when i went offline- not anymore i just downloaded it and it works great.
Google Gears is very exciting and addresses a lot of issues that tools like Google Documents have been criticized for not allowing offline access to. Scoble has a video of Google's Brad Taylor talking about the release, at the end of the video he says that a developer came up with this during his trips on the Google bus from SF to Mountain View- wait don't that have WIFI on those buses?
UPDATED: i figured video was not going to be available but i just noticed offline gets no images-graphics are important on many blogs. :-(

Friday, March 16, 2007

Google acquires Gapminder- great when can i have it for creating client presentations?


Yet another acquisition announced today by Google. This time Gapminder's Trendalyzer software.

"Make sense of the world by having fun with statistics" the Gapminder site tells us- ok i am in and want to have some fun.

I just watched the TED 2006 presentation by Hans Rosling founder of Gapminer and could imagine the Google people in the audience thinking hey- how do we get this on board with our strategy. During the presentation Rosling uses the Gapminder software to present the developing world including interesting data on an economic, heath and political perspective- all great but it is the visualization tools that is driving the presentation and that is why Google just snapped them up.

During the video Rosling calls for the end of boring statics- alleluia! and for linking data to design- double alleluia!!

Rosling sees a lot happening in data in the next few years- yes thanks to this Google acquisition i see it also- taking the data visualization tools and delivering it to the end user not power users- just like what they did to the search industry.

In the video he says that students and policy makers get very excited when they see the promise of data visualization? hey you know what, i am very excited about a potential tool that as a corporate user i can use to tell a story.


So what i would like to know?


  • How hard is it to create these visualizations? do they even know how long i spend doing slide decks to not even get to 1% of the powerful story Rosling is presenting in this video. Oh wait- wasn't there some recent buzz about a Google PowerPoint Clone coming soon?

Marisa Mayer says more information will be made available soon. (On a aside-i will be at BlogHer next week and hope to finally meet her).

Thursday, March 01, 2007

Web 2.0 in the Enteprise and What Google is doing about it- but what about adoption?

I went over to the InformationWeek website to read the article about how most business Tech Pros are wary about Web 2.0 Tools in th Enteprise and then ran into this article from Reuters annoucing a partnership between IBM and Google that will allow IBM WebSphere users to choose from 4,000 existing Google Gadgets services -- mini-Web applications that adminstrators and/or users can add with the click of a button onto public sites or internal office intranets/portals. This move i believe will lead to a much bigger use of Google Gadgets in enterprise environments that are traditionally resistant to the use of such applications and others will follow- an interesting question of course is how Microsoft's SharePoint services will follow.

From the first article on tech pros being wary of adoption Web 2.0 the following issues were raised:

  • concerned about security > Well the new IBM implementation of Google Aps will manage the security issues, and the use of the gadgets can be audited and tracked. When things are build into the software there is a much easier business case to make


  • return on investment > Well there are many types of ROI statements one can make from productivity gains to increased sales- but the bottom line is how much is this going to cost me and what are the returns on that cost- well with this anoucement Google Gadget features are available at no cost to companies who have purchased WebSphere Portal Version 6.0 and also customers of WebSphere Portal Express


  • Concern about staffs' skills in implementing and integrating new Web tools > seems like plug and play with the Google integration into IBM- and it even goes beyond easy install for the technical folks- The end user decides: We no longer need to go off and call a technician," Larry Bowden, vice president of the IBM Lotus division for portals and Web services said. "The power has been turned over to the people who know best. You know best."

Collections of easy-to-install widgets, portlets, web parts etc. are not new to enterprise portal software suites- they are standard built in offerings. i have been integrating content into portals for many years and we have had partnerships with all the big portals- IBM Websphere, Microsoft SharePoint, Oracle, Plumtree, BEA etc. Although we have had successful implementations- it seems that in the long run, all of them still suffered from one big issue- user adoption.

I agree that companies like the ones mentioned in the article on Web 2.0 "shouldn't bet on employees flocking to these tools without a push". Adoption campaigns are key and in today's world you need to find the early adopters and make them the mouthpiece of what you are trying to achieve. The article mentions that Procter & Gamble is running an internal marketing campaign with the tagline "connect, converse, accelerate" as it rolls out real-time communications, a collaborative content portal, and desktop search. According to the article, Steve Ellis, executive VP of Wells Fargo's wholesale solutions group says that the Wells Fargo, IT and business departments are working together to develop only those applications employees need most and it seems that they have management support all the way because they are not 'bothering to cook up a dollar value for each collaboration app'. "I can just go out and tell our boss I know we'll be better off."- Lucky them would be great to learn more about how they got to that stage- or is their management just more enlightened?


Sunday, February 18, 2007

Google again harnessing the power of community this time for translations

Google's new feature for their Google translate service is harnessing the power of a community of users to make their translation services better (via Techmeme). Similar to the Google Image Labeler which is a 'game' that pairs you up with another player and allows the community to work together to add image tags- this new service allows users to suggest better translations from the ones that the automated translator offers.

It seems very easy to use per the screen below of my blog translated into Russian- but i could not find any information about what happens once a user submits a suggestion for a better translation. I would suspect that they have to have some sort of quality process before they take one translation over another- and just like in Google Image Labeler when the tags are attached to images only when two 'players' use the same tag (a match)- i would suspect that multiple people would need to match that translation identically before it gets fed into the computers for making the translation software better at translating web pages.


Automated translation and assigning 'aboutness' to content via categorization in multiple languages was a topic that came up during last week's Robert Scoble's ScobleShow interview with Clare Hart because of our handling of content in 22 languages (video to be made available hopefully soon on Podtech). The conversation was around the monitoring and measuring of new media (e.g. blogs, podcasts, vlogs) for Corporate PR, Marketing and Product Development professionals and how with the day over day growth of non-english content, teams that do not have multiple languages on staff still need to be able to monitor what consumers are saying in their local language.

It is interesting to see the languages that are in BETA for this service because it seems to be indicative of the languages of growth on the web and in the social networking world- which i am sure Google hopes will make their translation services for those languages better through participation by those actual communities. They are also obviously the 'hardest' of the languages to do computer translation due to their non-Latin base.