Yahoo Inc. plans to begin scanning books and collecting other media content in an online database rivaling Google Inc.’s efforts, according to a media report.
An unusual alliance of corporations, nonprofit groups and universities have announced an ambitious plan to digitize hundreds of thousands of books over the next several years and put them on the Internet, with the full text accessible to anyone.
If we get this right so enough people want to participate in droves, we can have an interoperable, circulating library that is not only searchable on Yahoo but other search engines and downloadable on handhelds, even iPods, said Brewster Kahle, founder of the Internet Archive.
The project, to be run by the newly formed Open Content Alliance OCA, was designed to skirt copyright concerns that have plagued Google’s Print Library Project.
The effort is being led by Yahoo, which appears to be taking direct aim at a similar project announced by its archrival, Google, whose own program to create searchable digital copies of entire collections at leading research libraries has run into a series of challenges since it was announced nine months ago.
The Authors Guild sued Google last week, alleging its scanning and digitizing of copyright protected books infringes copyright, even if only small excerpts are displayed in search results as Google plans. Google argues that the project adheres to the fair use doctrine under U.S. copyright law, which allows excerpts in book reviews and the like.
Unlike Google, Yahoo will scan and digitize only texts in the public domain, except where the copyright holder has expressly given permission. The OCA project also will make the index of digitized works searchable by any Web search engine. Because Google is restricting public access to excerpts of copyright protected books, it is maintaining control over the searching of all the digitized texts in its program.
The Open Content Alliance’s scale, at least initially, is more modest. Yahoo, of Sunnyvale, Calif., is funding the scanning of 18,000 books identified by the University of California system as being part of the American canon.
The consortium’s other partners are Adobe Systems Inc., Hewlett-Packard Co., Prelinger Archives of San Francisco, and O’Reilly Media Inc., of Sebastopol, Calif.; the University of California and the University of Toronto; the National Archives of the U.K. and the European Archive in Europe.
The University of California’s 10 campus libraries have about 33 million volumes, of which an estimated 15 percent are in the public domain, said Daniel Greenstein, associate vice provost and University Librarian of the California Digital Library.
There is good evidence to suggest that if people see that a book is out there, they will buy it. Print sales either increase or are unchanged, he said. We have not once seen data to suggest that open access, at least to published printed works, decreases sales.
Greenstein said that contrary to publisher concerns that people will choose not to buy books if they can read or download them free online, the ability to easily find books on the Internet will broaden the public’s exposure to them and is likely to increase, not decrease, sales.
By exposing more people to scholarly works, the OCA project could contribute to improved research and help reverse the trend among publishers of cutting back the number and print runs of books, said Lawrence Pitts, chairman of the University of California Academic Counsel Special Committee on Scholarly Communication.
The University of California Press is likely to participate in the project, said Lynne Withey, director of the UC Press. I’m all in favor of extending the availability of both books and journals in digital formats, she said. So anything that does that in a way that respects authors’ copyrights and also allows publishers to stay in business is a good thing.
The OCA is appealing to publishers and other libraries, universities and archives worldwide to offer materials as well. This is an international effort, not just domestic, said Dave Mandelbrot, Yahoo’s vice president of search content. For example, we would be very eager to integrate French content into the Open Content Alliance and are working with people in France to make that happen.
After Google announced its effort, the French government said it would embark on its own book digitization project, complaining that the Google plan would only accelerate the domination of the English language over other languages.
The OCA effort was applauded by publisher and author groups who have been critical of Google’s effort, including the Association of Learned and professional Society Publishers, the Text and Academic Authors Association, or TAAA, and the Authors Guild.
The OCA also is looking for ways to help publishers be compensated for offering copyright protected books to the repository, said Mandelbrot. We are working directly with publishers to come up with business models to encourage them to come up with ways to make works publicly available, he said.
Consumers will be able to search the contents of the Open Content Alliance’s database and download the entire content of any work, such as a scanned copy of a book.
When asked to comment on the Yahoo project, Google spokesman Nate Tyler said, we welcome efforts to make information accessible to the world.
