Tuesday, 26 March 2013

Machine readable rights - a new world of frictionless licensing?

There was a distinct buzz as representatives from the news industry, image libraries, publishing, software, photography and the legal profession gathered in Amsterdam for the Machine Readable Rights and the News Industry: Opportunities, Standards and Challenges,  which formed part of the IPTC Spring Conference. The pot at the end of the rainbow is  'frictionless licensing', a vision of licensing transactions on a seamless automated workflow between content owner, publisher and end user.

Although the IPTC is rooted in the news industry, its scope is broader, and the issues discussed at this conference were widely seen as critical for all sectors of the creative world, and touch on issues of copyright, business viability and the future of all content licensing across media sectors and boundaries.

As an organisation representing the image licensing industry CEPIC is working to secure future business for image creators and image libraries. CEPIC’s participation in the European RDI (Rights Data Interchange) project was described by Sylvie Fodor. She outlined the need for unique identifiers for images and for methods of reuniting images with their rightsholders in an environment where the proliferation of so-called orphan works is threatening the very business of image licensing. The RDI will provide a test bed for CEPIC as a registry allocating unique identifiers for images, and as an exchange, where users are directed to licensing outlets for images. Partners in the project are Getty Images/PicScout, AGE/THP, Album and PLUS.

Eugene Mopsik from the US photographers association ASMP also called for unique identifiers. It is easier to right click and download than to license an image, he said, and called for a professional attitude on behalf of image providers so that rights data is supplied along with the images. It is no longer good enough, he said, to wait for misuse and then litigate. Supporting the use of the PLUS standards, he reminded the conference that a Picscout survey found 80% of web uses were unauthorised. The logic here is - make it easier for users to do the right thing and license the image. Mopsik supports the use of the PLUS registry as part of the solution for orphan works, and indicated that linking photographers direct to a PLUS registry is a 'no brainer' technically. The hold ups are more political and administrative than technical.

Licensing deals between press agencies are a complex and time-consuming business, said George Galt, Associate General Counsel for Associated Press. He highlighted the fact that lawyer time could be radically reduced if there were standards in rights expression, citing an example where a 3 year contract between agencies was finally signed 6 months after it was due to expire. There is a long list of rights expression ripe for standardising; exclusivity, prohibitions, retained rights, payments terms and copyright are just a few.  Getting the key terms standardised would go a long way to improving workflow he indicated.

Making money from content is the key driver to standards and automation. Andrew Moger, Executive Director of News Media Coalition, previous picture editor at the Times newspaper, said that as far as he could see everyone has worked out how to make money out of photographs, except photographers. He talked about conflicts over image rights in the sports world, citing the recent lock out of Getty Images from the Australian Cricket tour of India, and consequent ban on distributing images in a protest supported by all agencies including those allowed at the event. The issue of exposure by the press is taken for granted, he said. Imagery is valuable to businesses of all kinds, and ways must be found to retain that value for the content producers.

Thomas Hoppner, from law firm Olswang raised another issue which should concern all media sectors - the power of aggregators like Google to control and skew the market.  The Google site has become a place where people read the news online, and content is gathered from news sites, placing the search engine giant in direct competition with news providers. Content is supposedly protected against copyright infringement by the Robots Exclusion Protocol, but in practice if the protocol is used to block search engine access to material on a  news site, overall findability is severely compromised, making the site more or less invisible on the internet. Such is the power of Google. Negotiations are underway, but the balance of power is not equal, and there is a danger that the content creators and providers will lose out to the aggregators.

The Newspaper Licensing Agency, NLA, works on behalf of major newspaper groups to collect revenues from schools, cutting agencies, and publisher feeds. Faisal Shahabuddin, Head of Product and Service Management at the agency pointed out that progress in the area of machine readable rights will depend on the bottom line. Will developments increase revenue or reduce costs, he asked, leaving the answer open.

Returning to the issue of copyright, Michael Steidl, Managing Director of IPTC, described the recent IPTC study of metadata retention in images placed on social networking sites. Millions of images are uploaded to social networking other sites daily, and the study showed that FlickR, Twitter and Facebook currently remove key copyright data embedded in their image by photographers. Google + and Tumblr do better. The results can be found at the Embedded Metadata Manifesto site, which was set up by IPTC to campaign on retaining metadata embedded in image files. This issue of removing embedded data gets to the heart of copyright protection. The technology is there, but software developers ad suppliers need to more aware of the need to retain embedded data, and critically, to implement the technology to enable that. In a world where software development is based on user need, the IPTC wants to make both software companies and their clients aware of the need for metadata to be retained. The words 'noone is asking for it' should not be allowed to pass their lips.

John McHugh picked up the copyright issue from the point of view of a working journalist. Photographing in Afghanistan and other dangerous locations, he is acutely aware of the emotional and financial investment he makes in the images he produces. Dissemination of photography is exploding, he said, but revenues are not following, and attribution for images is thin on the ground. The IPTC metadata offers copyright protection but when metadata is stripped the photographer is left high and dry. Sharing and stealing is far too easy, he said, and photographers cannot not simply opt out of the internet. His solution is to take a step back and make it possible to watermark images direct from a mobile phone.  Marksta is a new company he set up to create an app which creates custom watermarks visible on the image, and allows photographers to enter IPTC metadata direct from their mobile phones. It’s not only professional photographers who are affected by losing their accreditation and rights. Marksta is also driven by soccer mums wanting to retain control of images of their children on the internet. The broader point here is that if copyright disintegrates in the social and non professional sphere, hanging onto it in the professional sphere will be nigh impossible if an anti copyright, use-what-you-want-where-you-want culture wins out.

So what are the standards that can be applied to expressions of rights and how can they be used? Graham Bell, Chief Data Architect at Editeur, has been working in the field for many years. Editeur manages the ISBN standards for books and has developed a number of other standards in the publishing industry. He distinguished between machine readable rights expressions, and those which are machine actionable, the basis for automation. Editeur has developed ONYX standards, which operate between publisher and retailer. So far there are few implementations, and this is a criticism often levied against standards bodies  - why is no-one using this standard? The answer, as those working in the standards field know well, is that it is a long game. If standards take time to develop, implementation often takes longer. The real world of commerce eventually comes to recognise the advantages of using standards, and it seemed from this conference that the time for machine readable rights is on the horizon. The people and companies who develop standards always have to work ahead of the curve.

The PLUS Coalition has produced a set of standards developed for use in the image industry. Jeff Sedlik and Ray Gauss presented the work of the organisation which recognised the need back in 2004. The international non profit PLUS Coalition created a matrix of rights expressions for use in the stock image industry, and a registry for images, creators, licensors and licensees, to enable rights data to be held online. This addresses the critical issue of where data is held. Data embedded in images is important, but there is wide agreement now that important data needs to be stored in the cloud. The principle is for embedded data to point to an online database or registry where it can be changed dynamically and so is always up to date. There are a number of important issues about who holds the data, how accessible it is, and who has access to what data. Both PLUS and Editeur have been engaging these issues for a number of years, and the registries reflect that.  PLUS has developed rights expressions to cover parties, industries, scenarios, licensing modes, time frames and much more. Adoption, as always, depends critically on user demand and software development. It's important at this point to note that none of the standards bodies forsees a future where there is no human interaction at all. It's more a case of let’s automate and standardise where it saves time and money and where it makes sense.

IPTC has been hard at work creating a rights schema for the news industry Rights ML using the Open Rights Digital Language ODRL.  Stuart Myles, Director of Schema Standards at Associated Press and Right ML Lead at IPTC presented the principles behind Right ML - that it is publishing specific, supports todays restrictions, and buolds for the future. The news industry, he said, needs increased automation to cope with the fast moving digital environment and sophisticated publishing relationships.

Creative Commons grasped a key aspect of licensing and rights years ago when it set up simplified rights expressions and icons for images on the internet. The user of a digital image needs to be able to be alerted to the rights situation, and there needs to be a presumption that there are rights associated with the image. The fact that the Creative Commons baseline is  that images will be shared is not at issue here; it is more important that it is clear that rights belong ito the creator or copyright holder, and that licences are actively granted not just taken for granted.

The use of Creative Commons licences has increased enormously; in 2007 they had 29.5 million works and in 2013 252.6 million works under CC licence. The expressions in CC REL (CC Rights Expression Language) code is legal (readable by lawyers), human understandable, and machine readable. The fact that many commercial image licensing bodies do not use Creative Commons lies in the failure of the codes to sufficiently define what commercial and non commercial uses entail. This can lead to users thinking, for example, that use of an image in a charity annual report is a non-commercial use. Anyone involved in licensing images will recoil at this understanding, and work in this area would be a prerequisite for acceptance of this schema by the professional licensing industry. This is however an important step, to bridge the gap between propoenants of 'free' and those engaged in licensing activity. A definition of boundaries will be essential to happy coexistence.

Andrew Farrow Project Director of the Linked Content Coalition (LCC) outlined the ideas behind the RDI project as an exemplary implementation of LCC.  LCC aims to provide a technical infrastructure to enable interchange of rights expressions between different existing media schemas so that rights can be read and understood across all media in an automated way. The RDI project with its 16 partners from different media sectors, starts in May 2013 and its objective is to demonstrate the range of dataflows across the media supply chain and show that the data can be interoperable.

Andreas Gebhard from Getty Images chaired the first session on the need for machine readable rights. Getty, as owner of PIcScout has long understood the need for copyright protection, and for mechanisms to link users to licensors. ImageIRC from Picscout used with ImageExchange allows images registered with Picscout to be located and tracked on the internet using fingerprinting technology and links users with the means to license them.

Zhanfeng Yue from Beiijing Copyright Bank Tech Co Ltd presented another form of protection copyright, the Copyright Stamp, which creates a visible and resolvable icon in the image, so easing the path for those who want to use images, and protecting the copyright of creators.

Generally, the mix of embedded machine readable data, fingerprinting technology and visual search mechanisms form the bedrock of copyright protection in the web environment. At a time where copyright and rewards for content providers are under severe pressure, discussions at the conference show a real desire to use available technologies to support content creators into the future.



Presentations from the conference can be found at the IPTC web site.

Watch out for the IPTC Photometadata Conference at CEPIC in Barcelona in June this year. Many of the topics raised here will be discussed further.

Thursday, 11 October 2012

Do you trust your data?

WE CAN HELP YOU 
At Electric Lane we help people set up image management systems. Data handling is top of the list in the early stages, and legacy data can be one of the biggest headaches if you don't tackle it at the right time. If you do, things will go swimmingly!

What I am saying here applies to any data transfer, whether it's associated with images or not. Anyone  transferring data from one system to another should ensure that the data is clean, separated, and granular. And if a rule can be applied to data handling, its should be automated. It's the job of a machine, not a person. (We mortals have plenty to do.). Given that data is driving much of the new internet based business around the world, it pays to get yours in a good state.

DATA AUDIT
The first thing we do is to audit what's there.  Then we set up a set of rules to apply to the data so that it is exported in a fit state, and configured as our customers want it. We invariably find problems with legacy data.

DATA FIX
Here are some of the things we find and fix. We use our own script, which can be configured to do just about anything with data (as well as images - that's for another time).

Hidden characters
What you see is not always what you get. Hidden characters are often present, especially in exports from databases. These characters, which include character returns, tabs, and other non printing instructional characters can affect your workflow later on. A carriage return entered in a descriptive field in a database can result in the data being divided into two columns or rows in the export. Horrible!  We clean these out out with an invisible hand so the data is exported as intended.

Unseparated data
Often data which should be separated has been merged in one field, sometimes without separators.  In the art world for example, a field might read  Vincent Van Gogh, 1888, Vase with Sunflowers. There are 3 fields here which need separating in the database, and that's easy if commas are only used as separators. Even if there are no commas we can separate the data if there is a consistent pattern of different types of data. So if all the captions read in this order Vincent Van Gogh 1888 Vase with Sunflowers we could separate the data by applying a rule which places separators before and after the date.

Commas
One of the common data formats for delivery is CSV (comma separated values), with commas as separators. This means that other commas in your data can confuse the export. If you have descriptive text with commas, and separators are needed, it's better to export data as tab delimited format, otherwise your sentences will split at every comma.

Formatting issues
There are at least 5 different types of quote mark!  In a plain spreadsheet environment formatting symbols like quote marks, hyphens and spaces will be substituted by characters.  But the substitutions can be different in different environments so the answer is to clean them out before they get to the spreadsheet. (For example you may have seen this appear %20 when you paste a web address into your browser. %20 is a substitute for a character space.)

Diacritics
Diacritics (accents and non-roman alphabet) are not always correctly interpreted by software. They are often not allowed when keywording images, for this very reason. If appropriate, we can export data without diacritics, or we can ensure that it data is correctly exported.

If you are interested, you can see the possible variants here in the ascii table. (Unicode enables all characters to be correctly handled, but data needs handling differently and not all software is unicode compliant.)

Dates
Dates can be a horror. We know that dates entered in some earlier versions of some software (including Photoshop) can change when read in other versions.

In a spreadsheet, data in a cell can be string (characters), number, or date format. The spreadsheet will reformat your dates to the preference you have set. (Day-month-year, month-day-year, and so on) You may well have encountered this in Excel, but it applies to all spreadsheets. In XMP (the Adobe data model used for embedding data in images)  a date can be year only (no day or month). This kind of date is treated as a numeral by spreadsheets (just a number) so falls out of order with the date values when being sorted.

We have a way round this when we want to analyse data. We insert a string character so that when dates arrive in the spreadsheet cell they do not reformat. For import back into an image or into another database the appropriate date format needs to be applied.

For historical archives there is an additional problem with circa dates, which cannot be expressed in date format and require an extra text field.

Keywords

We sometimes find that databases export keywords with non standard separators which are not recognised by spreadsheets. Replacing spaces with comma separators causes problems for compound keywords, like Leather Jacket. We review the data before it gets to the spreadsheet to identify the hidden separators and substitute standard separators so that the keywords read correctly.

Mapping
You may want to map data from one field into a field of a different name. Along the way, you may want to split or even merge data for various purposes.  Everything is possible so long as the route is clear and no information is lost along the way. People have stopped talking about the all singing all dancing single data structure. There is recognition that legacy systems are here to stay in one form or another; that differents chemas need to be able to talk to each other.

 So these are the things we can sort out for any kind of data.

BRINGING IT ALL TOGETHER?
The problems of data are exercising the minds of a number of people around the world. The European database of art is one example where data from a number of sources is pulled onto one site for a search across all contributing sources. The data structure is based on Dublin Core, which is extensible but not supremely fitted to describe imagery (there are only 5 fields which map directly between IPTC and Dublin Core). So there are some inconsistencies turning up in the data fields displayed online.

But there are problems with these 5 fields too, if Dublin Core is not qualified. For example, IPTC now has two creator fields, one for creator of the photograph, and one for creator of the artwork.What can happen is demonstrated by an image I saw once on the Getty site. The Mona Liisa painting was  displayed online, and the photographer was listed as Leonardo Da Vinci. That was before we had an artwork creator field in IPTC.

Because we find outselves working at the interface between collections management and image DAM systems, we are getting more involved in collections management data structures. We have set up a Cultural Heritage metadata group for people working in this area, and hope to create a common set of fields  for heritage works from which to produce a subset of new fields for IPTC. Clearly we will not be reinventing the wheel, and are working with people involved in VRA Core, Linked Heritage, and others. I have also been involved in the ARROW PLUS project, where standardised querying of data is an important element in the effort trace rightsholders.  More on all of this later.......

See also my 10 rules for image metadata in the CEPIC/IPTC Image Metadata Handbook

If you need help with your data contact sarah@electriclane.co.uk .

Friday, 3 August 2012

Copyright Works!

After what must have been a very intense 8 months,  'Copyright Works - Streamlining Copyright for the Digital Age' by Richard Hooper and Ros Lynch has been published (I will call it the Hooper report). Commissioned by the government, this independent report is ambitious in scope, bringing into one document the thoughts and experience of a variety of creative industry sectors, all with their own particular drum to bang.

The main points relating to the picture industry will be listed at the end, and if you are acquainted with all the various strands of thought which have informed the report, go straight there. Here's the background for everyone else....

The big idea at the start was to look into the setting up of a digital content exchange. For many of us, the idea was vague. What exactly is a digital content exchange and whoever would be in charge of setting it up? In the course of discussions it became clear that the picture industry already operates digital content exchanges in the form of image libraries. Now there's a concept we understand - a user goes to a place online, asks for content, finds out what rights they can buy, pays the money, downloads the image. Easy in our industry, we've been doing it for years. So what's the need for change?

Licensing rights can be more complex in other industries like music and audio visual where multiple rights may exist in a work, and several, sometime overlapping, collecting societies are responsible for handling rights. For the user it's a dogs dinner, and there are attempts to streamline some of those databases into something as near as possible to a one stop shop for people who want to licence content.

The problem in the picture industry is a little different, as the Hooper report recognises. We have user friendly access for people who want to buy images, but there is a real problem with image identification. Images that escape from databases are mostly floating in the websphere without metadata, without an identification label, and with very limited means of finding rightsholders. They easily become orphaned, and metadata is almost routinely stripped from images when they are uploaded to the internet by anything other than the most professional image library software systems. This affects anyone who takes pictures, the amateur as much as the pro. If the concept of copyright is to be retained - and the Hoooper report recognises the importance of this -  there must be an effective way to label all images so that the information has a chance of sticking to the image.

For the general image user, things are just as frustrating. If you  find an image on the internet you want to use - whether for commercial use, for education, for a powerpoint -  you will have a hard time locating the rightsholder. You can use image recognition - Tineye and the Google's search by image - but what you often get are pages of results showing where images have been used, passed from site to site without a clear licence. How then to find the legal place to licence or use an image?

At the IPTC conference at CEPIC in London this year (see my last blog), we raised these issues. The position  I took, in my paper 'Orphan Works and Image Licensing' for the ARROW Plus project for CEPIC was that to be sure of identifying the source of an image you need  a verifiable identifier embedded in the image so that rights information can be resolved (by a url, or by a registry), and some kind of visual icon to tell you that such information exists.  Visual recognition can help identify uses of the image, and digital fingerprinting can embed an resolvable identifier into the pixel structure which is harder to remove than data embedded in XMP fields. (see presentation Avoiding Orphans)

Although image libraries have identifiers for images (picture numbers) which are unique to their organisation, what's missing is global uniqueness, which can be delivered by properly accredited registries. Other industries are ahead of us there, books have ISBN numbers and the music and AV industries have their own standard identifiers.

It's often assumed that if there is a standard numbering system, everyone has to change the way they work, renumber their images, change their database. This would cause chaos and confusion and put hard pressed businesses out of business. But there has been general recognition by people working on identifiers and registries that there is no value in making people change the way they work. The value we can add for the future is in linking what people are doing, and translating their efforts so that data becomes interoperable and useful in a wider sphere.

 Verifiable globally unique data can be added to images by a registry system. The IPTC schema has already make provision for registry data to sit alongside a picture library or photographers own numbering reference.

PLUS has thought all this through. I won't go into how it works here, except to say that like all workflow developments in the future, it will be down to the software people to automate workflows to make it easy for users. The thinking behind the PLUS schema is sound, and anyone who is serious about the future of the imaging business is advised to engage with the ideas they have developed.

Other problems for images? How can they be connected with other media so that people can license across media types. CEPIC is part of a proposed European project the RDI or Rights Data Integration project, which is being led by the Linked Content Coalition (LCC) and includes partners from various media sectors (It is expected that the project will be approved by the EU). The broad ideas is to set up a technical framework for allowing different rights schemas to talk to each other. Again, this is not a matter of trying to squeeze all sectors into the same way of working, but rather a way of setting up a mapping process which can be automated, so that systems can speak to each other with a defined vocabulary at the core. Its a little like Esperanto, which was designed to be a bridge between languages, to facilitate communication (and peace) between nations.  As with language translation, we know that the data mapping process can never be perfect, just good enough. CEPIC will be a content exchange partner, testing the concept of  a European Content Exchange for visual works.

The Hooper Report reflects the fact that those involved have listened carefully to our industry (as well as to all the others) and has not fallen into the trap of treating all sectors in the same way. The fascinating part of it is that the work leading up to the report has actually facilitated cross sector engagement and made people in the creative industries think ahead.

So here are the main points relevant to the image industry:

Metadata stripping
The report recognises the problem of metadata stripping (P15 of the report). This is significant. There a call for image using organisations to develop a voluntary code of practice committing not to strip metadata from images and not to use images from which metadata has been stripped. Further, that software developers work with the image industry to find a solution enabling images to retain metadata when posted to the web. (This is technically not difficult. The will of the software purchasers is a key element in shifting software into the modern age.) And the report recommends that the government work with our industry as far as practicable to help find a solution to metadata stripping.
We can all be rightly sceptical of voluntary codes of conduct, but this is very good news. A great boost for the  IPTC and the EMM (embedded metadata manifesto)!

Social networking sites
The report mentions social networking sites like YouTube and FlickR (P73)and the need for capturing data as content enters these systems. The IPTC has been investigating metadata stripping in social networking sites, and we see this as a significant barrier to copyright protection. The fact that the Hooper report recognises the importance of copyright for pros and non pros alike is very welcome indeed.

Linked Content Coalition
(LCC)
The work of the LCC  is explicitly supported by the report. This is good for data interoperability, and for the proposed RDI European project in which CEPIC will play a part.

The Digital Copyright Hub
The hub is conceived not only as a DCE (Digital Content Exchange) although it will function as such in some areas. Rather it is a way of linking different exchanges and registries. Bringing together the best of technology already being developed, the Hub will be a not for profit entity, drawing on the experience of the Copyright Clearance Centre in the US, the Linked Content Coalition in the UK, and the Technology Strategy Board's recent work on a Digital Licensing Framework (DLF)- a project in which National Maritime Museum, Tate Images, V& A and Pearson have played a part.
The Hub, it would seem, could be all things to all people - a linking mechanism for DCE's (read image libraries in our sector) and marketplace for photographers, a place where rights can be queried across media types, an orphan works registry, and part of a diligent search.  The problems are more likely to lie in governance and technology creep (technology on the ground not keeping up with opportunities) than in the technology itself.

This is the opportunity that the Hooper report has given the creative industries in the UK -  to be at the forefront of linking technologies and user friendly rights exchange hubs. The detail to be grappled with is staggering, but someone had to get outside the media silos and understand the opportunities technology is offering, and that effort could present a turning point for creative industries.

Perhaps we are all a little light headed in the UK at the moment (those of us who haven't escaped abroad). After Danny Boyle planted a flag for creativity and technology at the Olympics opening we all felt a stirring of pride for what we can do. After Wiggins and our various gold medals our feet have perhaps left the ground. But it's hard not to get excited when a government commissioned report backs the out-front ideas we have been playing with over the last few years. It's good to be heard, and now the hard work begins.

Ros Lynch will have her work cut out for her in the first year of spawning the industry funded Digital Hub. There will almost certainly be issues for smaller and poorer industries playing with the big boys. But for now we can be glad that Richard Hooper and Ros Lynch have so intelligently pulled together the ideas and aspirations of so many differing interests and come up with a roadmap for copyright protection and licensing activity which, in its outline, makes sense to so many people.

Sarah Saunders
4 August 2012

Monday, 11 June 2012

Find the Rights - IPTC Conference 2012

It was great to have the IPTC Photo Metadata Conference in London this year. Metadata has moved now to a central position in the image licensing industry. Our topic covered two aspects of search-  finding the image and finding the rightsholders - both essential for the future of the image business.
 The full agenda for the conference is shown on the IPTC web site, and I report here on the sessions I was closely involved in. Overall, the conference was a great success, with excellent speakers and very good feedback from attendees. Image search was a major topic, with sessions chaired by David Riecks including reports on visual search techniques, controlled vocabularies, crowd sourcing,  linked data and indexing images for web search.

In the morning session 'Search  - finding the rights', which was moderated by Linda Royles, I gave an overview of ways of finding and protecting image rights.  This encapsulated the thinking in my paper 'Orphan Works and Image Licensing' written for CEPIC as part of the ARROW project. The presentation (not long and very visual) can be downloaded here. The main points I set out were along these lines:

1. Positive identification not orphan grabs
Positive identification is the opposite of the orphan works database concept. It encourages people to use images which have permissions, not those which don't. Methods include embedded metadata, unique identifiers linked to registries, visual search methods, and digital fingerprinting.

Orphan works databases may be needed in the heritage sector to enable institutions to digitise older orphaned works in their collections, but this is not a solution for currently circulating digital files which may have been orphaned because technology and behaviour have not yet caught up with the medium.

2. Labels to help image users
People need easy access to images they can legally license or use.  For this, we need identifiable icons on thumbnails and easy click through to source web sites. (see PicScout Image Exchange )

3. Copyright protection for all works
If copyright protection is eroded for images belonging to the general public, that will impact the creative industries as well. The guiding principle should be that permission is required for use of any image or artwork. It's not good enough to just use any old image because there's no information.  Creative Commons licences will become increasingly important, especially for non professional photographers (and remember we are all photographers now!) but the term 'commercial use' needs to be properly defined, so that image licensing opportunities are not subsumed by the term.

4. We need registries and content exchanges
An image with no metadata should not be released from copyright, but we do need to change the way we work to secure digital images for the future. Digital content exchanges and registries will probably be needed to secure the rights of creators and to promote licensing and easy access for users. Identifiers require registration bodies in order to be authenticated,  global identifiers require a global network, and users need easy access via a user interface.

Content exchanges  consist of access to identified content, rights information and delivery mechanisms. They are being promoted as a way of finding rightsholders in a cross media environment. In the UK, The Intellectual Property Office (IPO) will be reporting soon on its consultation about a UK Digital Content Exchange (DCE)

Image libraries are in fact already content exchanges. In some ways they are ahead of other industries. Where the image library industry needs development is in the area of unique identifiers, so that images can be properly tracked in the web environment. The PLUS registry answers many of these requirements including that of a network of registries to form a global registry.

Other contributors in morning rights session
Antoinette Graves from the Intellectual Property Office (IPO) outlined the framework for the IPO  consultation on copyright, remarking that the major area of concern is the heritage sector where digitisation projects are held up by orphan works and the resulting legal uncertainties for the heritage sector. The consultation is ongoing and results will be published later this year.

Mark Bide  from EditEUR argued that investment in content made by creators is now lagging behind the returns gained in the internet environment. It now seem easier to profit from other people's creativity, and it is important to turn the tide and make investment in content pay for the creator, as it should. The answer to the machine, he argued, is in the machine.  Technology and automation can bring about a revolution in licensing while preserving copyright and benefiting content creators in all media sectors.

 Nancy Wolff  from New York law firm CDAS warned that orphan works legislation is coming. DShe believes it is critically important for the picture industry to get to grips with issues and solutions now.

Afternoon Session on Orphan Works  
I chaired a break out session which pulled together some of the topics raised in the morning, and allowed for discussion between the people involved.


Offir Gutelzon  from Picscout demonstrated how the Picscout registry provides tracking and identification services for image suppliers and help for image users to access licensable images from an internet image search.

Cathy Aron, Executive Director of  image library association PACA described how  associations in the US are combining to reach agreement on issues relating to orphan works, in the expectation that orphan works legislation (which failed to enter the stature books in the US last time round) will eventually be passed. Issues like diligent search, the need for registries and the concept of safe harbour for cultural institutions are on the table, and it was agreed that CEPIC and PACA should stay in touch on orphan works related issues.

Paola Mazzucchi from EU project ARROW demonstrated how the system works to link and access data about rightsholders relating to orphaned books.  She  discussed the difficulties in accessing rights data about images which are not credited, and talked about the feasibility study currently underway at ARROW which will look more closely at easy of finding image rightsholders.

With legislation pending on both sides of the atlantic the urgency for the image business to find orphan works solutions  was stressed on all sides -  it was recognised that this is a global problem which will need global solutions.  An important step would be to find common ground on the essential elements of a diligent search.

Monday, 25 October 2010

The known and the unknown - keywording for visibility

Why is everyone talking about keywording? People in the image industry are scratching their heads about ways to keyword their images. Now the web is buzzing with visual material, words need to be deployed intelligently to ensure images can be found by the people who need them.

Technology offers other ways of finding images, you may say, which don't require so much human input. Visual recognition techniques do offer clever ways to look for images, but computers can only learn from the way humans keyword the images in the first place, and they are not very clever at understanding abstract concepts. How do you explain to a computer all the different ways of expressing the idea of freedom, for example? Can love only be expressed by the shape of a heart, or a smile between two people? Human interpretation is still needed, and computers are still taking baby steps at recognising 'things in the picture' like trees and tables, never mind the more abstract and subtle signifiers found in visual material. So what we are looking at, for some time to come, is human tagging of images to make them findable.

The problem anyone keywording images faces is this. Language is a wonderful, expressive tool for communication, there are many ways to say the same thing, and words often have more than one meaning. The word I use to tag my image may not the the one used by the person searching for it. They may use the plural where I used the singular. They may use a different spelling, or different versions of a language like American and UK English. And then there are the requirements of multi language searches.

Any good tagging system needs to scoop up all the variations of a word so that whatever word the searcher uses, they will find their way to the image. Words need to be uniquely defined, so you can tell the difference, for example, between orange the colour, and orange, the fruit. There may be broader terms than the one you first thought of which may be useful, so your image of a train should also appear under a search for transport.

The way to achieve consistency, and to scoop up all the appropriate terms, is to create a controlled vocabulary. With a set of preferred terms, and their synonyms, the vocabulary is usually structured in hierarchical way to include broader and narrower terms. Vocabularies for use with images vocabularies have been informed by work done on text search in the library sector, but they have developed further to include concepts specific to visual material. One of the big advantages is that properly controlled vocabularies can be translated - just once- so that searches can be made in different languages.

Can a single vocabulary describe the entire world, the universe, and everything in it? Yes, if it has top level terms broad enough to cover everything, a logical structure, and sufficient depth to reach down to a granular level.

How does it help in practice? The vocabulary is embedded in the software both at the keyboarding and the search stage, creating automatic links between words and effectively automating much of the keyboarding effort. The keyboarding operative, with a well designed CV and good software can concentrate on interpreting the image for the user audience. Thats the part the machine can't do.

People in the stock image industry have been working on this for decades, and have come up with some pretty good systems for keyboarding, led by teams in large agencies like Getty and Corbis. Now it's time for everyone else to sign up for productive and accurate keywording, learning, where possible, from experience already gained in the industry on keyboarding and customer behaviour. The benefits will be felt not only by smaller picture agencies and photographers, but also in the wider world. Imagery is playing an ever greater part in company DAM systems, where the level and quality of retrieval makes sense of investment in this area. A picture may be worth 1000 words, but without words, a picture may be lost forever.

At Electric Lane we have been increasingly involved in creating vocabularies for image collections. We are also working with the standards body IPTC on a project to create a standard vocabulary to help collections of all sizes raise their keywording standards and make their data more interoperable.

We are offering a one day course, Keywording, on December 7 in London, run by Electric Lane Associate Liisa Kaakinen, a stock image industry keywording and controlled vocabulary expert. The course covers professional keywording techniques and the vocabularies that lie behind them, applied to still and moving images. For those wondering what to do about keywording, this session provides an essential step to understanding the process, the gains, the resources needed, and how to maximise productivity.

For further enquiries about course content contact sarah@electriclane.co.uk, tel 020 7607 1415.

See also
Is Language a moving target
Multilingual Keywording
IPTC Mirror on IPTC Controlled Vocabulary Initiative
Google is not Perfect, Fran Alexander

Wednesday, 6 October 2010

The Semantic Web, Linked Data, and Knowledge Organisation

It's a journey of possibilities, the semantic web, and its one we're all engaged in in one form or another. I've been eyeing up the topic for some time, and the third UK ISKO Conference gave me the push I needed to look deeper into what the semantic web and linked data hold for the future.
Last time I was at the ISKO conference, in 2008, many of us there were baffled by the possibilities of linked data. If all this data was to be linked on the web, who would put it out there? Who would put the money up to produce the data? This year, answers to some of those questions emerged. It would appear that linked data is at the point of lift-off, and it's already being used in ways we can now understand.

I start by looking at some of the ideas behind linked data, and then follow the presentations at the conference, all of which clarified some aspect of the subject and gave us a glimpse of how it can benefit 'the rest of us', the users.

The Significance of Linked Data

Linked Data is part of the Semantic Web, and for those wondering exactly what that is, here's a quick explanation. Semantic Web is the term used to describe a Web environment in which the meaning (or semantics) of information is made explicit and therefore machine readable. Our brains can handle very complex information. When we say we want an apple, we can work out from the circumstances that we want the eating type of apple, not the company Apple or the Big Apple, or an Apple computer. Machines need much more explicit instructions to contextualise the exact meaning of a word. We know that by building up a series of simple instructions, starting with the basic 0 or 1 choice, computers can perform very complex tasks. The Semantic Web, a term coined by Tim Berners-Lee the creator of the World Wide Web, describes an environment in which information can be accessed and processed automatically in an intelligent way. Linked data makes sense of the Semantic Web by providing a framework for a network of related information.

We are accustomed to the idea of HTLM documents located at URL's (Universal Resource Locators) and linked together by hyperlinks. In the same way smaller bits of information can be assigned URI's (Universal Resource Identifiers), and these bits of data can be linked using RDF technologies. So we can take the idea of 'Cat' and assign it a URI. We can give another URI to a particular cat, say Dick Whittington's cat. Then we can link the two to increase the amount of information available.

To link data in the public sphere, it has to be freely available on the web. The idea of releasing stuff for free is becoming more accepted with even parts of the creative industries starting to looking for new ways to make money in an environment where the market price of media increasingly parts company from the cost of creating it.

Services may come to be the new currency. If you share data wisely, you can attract value to your company and its services, increase the traffic to your web site and gain an element of trust which is working capital of a sort. By releasing data for free you can signpost things you can place a pound sign next to.

The linked data community is part of the open source community and while many of us working in media have been struggling with the recession and changes to our industries, there has been a quiet shifting of the scenery in the background.

The information revolution is being powered by new ideas and by the growth in mobile technology. People are already buying and using Apps to entertain, inform, and to find their way around the world. A steady stream of up to date information plays an essential role . The processing of that data into useful formats and apps will spawn the businesses of the future.

Share the data and the apps will follow
The keynote speech at the conference was made by Professor Nigel Shadbolt, from University of Southampton, Director of the Web Science Trust, and the Web Foundation. Together with inventor of the World Wide Web, Tim Berners-Lee, he was a key figure behind Data.gov.uk, a UK Government project set up in 2009 to make official data available on the internet for anyone to re-use. The thrust of his presentation was that once data is out on the Web, other people can do things with it, and this opens up opportunities of benefit to both individual citizens and to the businesses who create new services using the data. Publish the data, said Shadbolt, and the applications will flow.

The British Government has more that 4,000 datasets, and the aim is to make much of this data public. Already on the data.gov site there are apps using data previously locked away in government departments. One example is the data on cycle accident location. Once that data is pubic, applications can be created to help cyclists avoid the accident blackspots. By sharing data, Professor Shadbolt said, you can bring eyeballs and brainpower to a problem.

The public can do little with endless sheets of raw data, but data can be transformed into useful applications, and that's the basis for a new kind of business . In a world of iphone apps and mobile media, that information can be just what you need when you're looking for a bustop or running for a train, or finding the nearest dentist, once it's make accessible.

More data is now being made available. Postcode data was once copyrighted but is now freely available. From January 2010 every local council has to publish all spending over £500. There is a new appetite for open data. Public data provides ways of holding public services to account, and Professor Shadbolt's view is that it should be published quickly, and it should be linked.

Government departments can profit internally from linked data as well, he said, gaining better access to their own data and making better sense of related data from other departments.

He sees linked data leading to more accountability, more localism, more arguments. But how do we know we can trust the data and the interpretation of it? Shadbolt thinks there will be a flight to quality. Perhaps there will be a sifting process similar to that on the internet where some data sets are more highly regarded than others.

A language for linking
Antoine Isaac is scientific coordinator for Europeana and researcher at the Vrije Universiteit Amsterdam. He works with Semantic Web technology in the cultural heritage environment, focussing on interoperability of collections and their vocabularies. He was involved in the design of SKOS (Simple Knowledge Organisation System) the language designed to represent structured vocabularies (thesauri, taxonomies and classification schemes) so they can be used in the Semantic Web environment.

Isaac gave an overview of SKOS, which is expressed in RDF and enables linking between Knowledge Organisation Systems (KOS). SKOS, he said, presents a way of expressing structured information so that connections can be made between different knowledge systems like vocabularies and taxonomies. Unlike OWL, the language for expressing ontologies (which are generally held to be more complex broader taxonomies with more formal relationships between terms), SKOS is designed, as its name implies, to be easier to use and less manpower resource hungry than OWL.

RDF (Resource Description Framework) is a way of making statements about resources (particularly on the web). These statements are in the form: Subject (John); Predicate (has the age); Object (20 years). This is what is called a triple.

SKOS is a form of glue which allows different classification systems to be linked in the Semantic Web environment. The benefits are the re-use and sharing of information and the linking of concepts.

Isaac described the steps to take: put your data on the web; make it available as structured data; use open standard formats (XML, RDF); useURI's to locate the data; link it to other data.

Evangelising Linked Data
Richard Wallis has been with technology company Talis for 20 years and calls himself a Technology Evangelist. Talis started in the library sphere based at the University of Birmingham 40 years ago and is now one of the leading Semantic Web Technology companies.

Talis offers training and applications of linked data for a variety of commercial and governmental bodies and is involved in data.gov. Linked data is being used by Walmart, Tesco, The Library of Congress, the BBC, The Ministry of Defence and many others. Talis helped the BBC to link its own data to other sources of data on the Web. The BBC Wildlife site for example links to background information DBPedia, the linked form of Wikipedia.

One of the opportunities in using linked data is to create mashups, which are ways of combining data from a number of sources to create a new service or resource. Talis has a good example of mashups on its web site, where Guardian data about politicians was linked to BBC data about programming. The result was a timeline showing the exposure of individual politicians in the Guardian and on the BBC on different dates. This could be extended to other sources to create a very rich view of politicians' media exposure, demonstrating the opportunities for presenting and interpreting data once it's out there.

Connecting Communities
Steve Dale calls himself a community and collaboration ecologist. He blends technology solutions with an understanding of how people can be encouraged to organise and collaborate creatively in a sustainable knowledge ecosystem. He led the project to create a community of practice platform for the UK local government sector.

Data is everywhere, he said, and people are faced with the question of where to go for information, which networks to join, in an environment where there is almost no connectivity, and data is hidden behind applications.

Local government is awash with data, but is it being used to it's full extent? How for example do you compare performance with other areas? What is the relationship between national indicators? Dale charted the Knowledge Hub which should create the links to provide the answers. Integral parts of this Hub are data mark up and search facilities, data integration and aggregation, forms based data entry for benchmark comparisons, public datasets, mashups , and Apps. Data attracts value from contributions by other data producers as well as technical and user communities.

It's a big change, he said and its about open architecture, open source, open everything. Addressing the issue of reliability - how do you know if the interpretation of the data is correct or whether it is misleading- he agreed that this will be one of the pain points going forward, responding to the point raised that data is collected in different ways by different local authorities. Maybe the fact that the public is testing the data through available apps will do something to increase public awareness of the dangers in data interpretation. As Dale said, perhaps tomorrow will be a statisticians world.

Finding partners in trade
Martin Hepp is professor of general management and e-Business at Universitat der Bundeswehr in Munich. He looked at the costs in GDP terms of keeping commercial markets alive, citing a 1920 list of just over 5000 types of goods, compared to the current market environment where its easy to count just 30 types of bread alone. Finding partners to trade with is what drives business, he said, and the search space expands constantly. To find a specific item is the aim of the game, and the internet makes searches easier, but we still spend a lot of time every day looking for things. The world wide web, he said, is currently like a giant shredder, ingesting structured data and spewing it out as unstructured text, destroying or shredding information we had at the outset. What we need is to retain data structure, link data elements by meaning and reduce the look-up effort required.

Hepp says the quality of vocabularies will define how easily data can be re-used. Taxonomists everywhere can rejoice. Their calling in not only not going out of style, it is becoming becomes ever more relevant.

For the last 8 - 10 years, Hepp has been working on a web ontology for e commerce called Good Relations. Data levels need to be sufficient, he said, for rule based transformation. For example, the product needs to be distinguished from the offer, the store from the business entity, the product from the product model. If you are buying a car, the product data may be registration date, condition, mileage etc. A different set of information, can be extracted from standard model data. Major businesses have seen the opportunities and have implemented the vocabulary, which has the immediate effect of raising their google rating.

Dublin Core and Semantic Relativism
Andy Powell is Research Programme Director at Eduserv and has been active in the Dublin Core initiative since 1996, where he is now a member of the advisory board. He is active in standards activities relating to RDF, Dublin Core and other digital library projects.

He reviewed the history of Dublin Core, which started with 15 original metadata elements to describe web resources. Now there are around 60 properties and classes in a well curated vocabulary. Dublin Core started labeling with html metatags, which were later ignored because of Spam. It started with broad semantics, and 15 'fuzzy' buckets for data which were a hangover from library catalogue cards, and it was a record-centric model.

The challenge then was to change to the idea of strings of data, using RDF to express relationships. The open world view needs to be sold in, and there's the problem of URI's. Should a URI be a locator of a web page or an ID? For those of us struggling to understand the question, Powell introduced us to a useful Batman blog by Chris Gutteridge at Southampton University on the subject of modeling the world. When creating vocabularies we face the question of how to describe the entire world in unambiguous terms.This is amplified when you look at ontogologies on the web and the use of URI's to identify bits of data. See the blog for a tour into madness and back.

The conclusion? As ever, it points to the fact that we can only find approximate answers to questions. The web gives us the ability to approach things from a number of different angles to gain information. There's no point obsessing about absolute accuracy, as you are in danger of descending onto a Kafkaesque universe from which there is no point of exit. Enough said.

Ordnance Survey hits the streets
John Goodwin who works in the research department at Ordnance Survey (OS) was tasked with looking into the Semantic Web and its implications for OS. He started by constructing ontologies in the ontology language OWL to describe OS data, and is now responsible for linked data published by OS.

A number of natural hubs are starting to form in the linked data universe, and geographical terms are among them. (see this 2010 linked data diagram and compare it to this one from 2008).

OS is at the forefront of the geospatial linked data web, providing a trusted source for developers. Goodwin highlighted some of the problems and confusions with geo data. An area such as Hampshire can be represented as an administration area, or a county as used in the vernacular. Boundary lines can change for administration purposes while general usage of a geo term remains the same. There are issues around overlapping boundaries, touching boundaries and partial overlaps. It's hard to uniquely define an area, as anyone involved in geo term taxonomies will tell you.

The result of work done at the OS is that now you can not only find a map of anywhere in the UK on the OS site, but also, those interested in creating other applications can access OS linked data to work with.

Managing thesauri without knowing anything about Semantic Web
Andreas Blumauer is founder of Semantic Web Company, which focuses on technology consulting, the media economy and metadata management. He described the stage we desire, where the technical people have done their job, they give us a user interface, and the user has no need to look under the bonnet.

Blumauer gave us a quote from Dr Chris Welty ' It is not semantic which is new, it is the web that is new'.

He sees SKOS as having just the right level of complexity, and the ability to introduce Web 2.0 mechanisms to the web of data. Web 2.0 refers to changes in the way developers write applications for the web, enabling users to share, collaborate, and create content.

Blumauer characterised SKOS as a hand adding and retrieving linked data to and from the Cloud. The way his company's thesaurus management software Poolparty adds and retrieves data from the linked data web environment is shown in this Poolparty demo. The advantages to business are reduced costs of content management, increased automation of data handling, better search engine optimisation and access to new services and mashups (new applications made using data from a variety of sources).

Everything is a thing!
Bernard Vatant is an expert in data modeling, migration and interoperability of vocabularies. He has worked with the Publications Office of the European Union (EUROVOC vocabulary) and French National Library on the on the evolution and integration of RAMEAU for the Europeana project.

Vanant said that everything can be represented as a sign. People, products, devices, places, and concepts from vocabularies can be represented and connected. Everything, in other words, is a thing. He spoke about the semiotic triangle of meaning . He highlighted the differences between terms, concepts, and things and said they should all be first class citizens of the Semantic Web. His presentation contains rich thoughts on semantics and their use, and can be found here.

Go and play
There is something forfor all of us here. I hope you will all go out and play, look at some of the links on web sites like the BBC, or look out for apps using the data that's been made available.

It will all soon be second nature so we wont have to ask what SKOS is, or what a URI represents. Remember the Web when we thought it was just a way of skiving off to surf?

Links

ISKO Blogspot The conference encapsulated for SKOS by Fran Alxander
Data.gov
Government apps list
Talis Nodalites blog with musings on linked data and Pavlova and many other things
The Basics of Linked Data Tim Berners Lee - from the horse's mouth
Jenis musings Why linked Data for Government?
URI or URL?

Friday, 21 May 2010

Metadata and the Future of News

IPTC Business Meets Technology day Paris, April 2010

The IPTC evolved to represent the needs of the news business, and much of the Spring Conference this year revolved around the adoption of the IPTC's G2 standards for the news industry. The IPTC's Photometadata Working Group focusing on visual content, has responded to the needs of the photographic industry to make image search more focussed. The Controlled Vocabulary project was presented at the conference for the first time.

G2, which provides an XML-based standard for the exchange of news items, is well supported by large agencies, despite the fact that some of their customers are working with old systems and are not ready to change.

The news industry is looking to a future where news gathering is networked and syndicated, where news items are linked across media, and where outlets for news include mobile phones, social networking sites, and more personalised news delivery.

New business models involve aggregating content across publishers in the same way agencies aggregate content from their suppliers. Planning is needed to fund inter-organisational processes. The supplies within a network can then focus on their core business.

The networking and syndicating of news require interoperable metadata. There is a growing realisation that some sources of data, for example on events or on personalities and entities, are best shared. In the new world it is argued, it is pointless for everyone to gather the same data. This is where linked data comes in.

Fran Alexander
from the BBC highlighted the need for different sets of metadata in different stages of the workflow and in different departments of a large organisation like the BBC. Rather than insisting on standardisation across the organisation, she believes that mapping techniques produce better results, with standards risiing as departments increasingly work together.
The Linked Data community, which has emerged from the ideas around the Semantic Web, promotes the idea of sources of data scattered throughout the web, with unique string addresses on the web called uri's. These web addresses can link to other uri'’s to expand the available information on a subject or entity.

The principle is that by linking data sets on the web, the requirement to reproduce the same work of data gathering is reduced, and more data becomes available by a process of organic growth.

The uri's are simply web addresses for lists of information about an entity (a person, an organisation, a building, anything that has a name and is unique.) That information could include details of date of birth, height, weight, schooling, job history, any facts which can (or could theoretically) be checked.

Someone searching for information about Barack Obama for example, might find information on his first job after college which may reveal other facts about him which are of interest
UK company Talis demonstrated their own platform which is designed to build linked data. The Talis architecture is used by Government departments and other clients to pull together and link sources of information scattered around the country.

Talis suggested that the IPTC could be an ideal linking hub holding trusted data which can be linked to other uri's and other hubs on the web.

The Okkam Project is a European funded project which has already assigned 7.5 million unique identities for entities and has created 16 applications for their use. Okkam is running a pilot project with news agency ANSA, which allows the agency to create an enriched newsfeed using News ML standards, which can hold links to other sources of information. Okkam works on the principle of a 14 th century philosopher who said we should not' multipy entities beyond necessity.'

The Okkam system is designed to be an open neutral system, which will be run b y the Okkam Trust once the EU project is over. The Trust would be funded by money from commercial applications developed by the project, and is keen for stakeholders like the IPTC to join the board of the Trust.

Okkam aims to fix reference names for entities, but recognises it is not the only company to do so. The aim is to create permanent identifiers, which is where it differs from projects like Open Calais, which is about entity extraction.

Underlying all this are considerations about the trustworthiness of data sets. The web site www.sameas.org finds entities which are supposedly the same, but not all the results are seen as sound. How rigorous the test for sameness will depend on its use.

The IPTC Photometadata Group Controlled vocabulary project was presented by Sarah Saunders from Electric Lane, who discussed the ambiguities inherent in keyword searches for images, and demonstrated the use of keyword predicates, a new feature of the proposed controlled vocabulary, which helps to reduce ambiguities. An example is the word orange, which can be used as a descriptive word, or as an object word. The word Paris can be used to describe the location of an image, a person (Paris Hilton) or a view of Paris.

The new IPTC keyword predicates will separate the differing uses of a word, and avoid unwanted search results. The draft vocabulary and the ideas behind it will be presented at the Photometadata Conference in Dublin in June.

More information from the Spring Conference in the IPTC Mirror.
Sarah Saunders
April 2010