Friday, January 30, 2015

Partial address validation can be easy, cheap and effective!

I was surprised to receive a package today (late ...) from a company in the UK which included in my address the province as NIERDEROSTERREICH (sic), a province of Austria, instead of NIEDERSACHSEN, the province of Germany where I live.

There's no purpose to including a province in a German (or Austrian) address - I'm flummoxed as to why so many German websites ask for it - but the UK company concerned has my province name (correctly) stored in their records. To me this looked like a software error, where the sender had been allowed to choose a province outside the country to which the package was being sent when entering details for the courier service. Parcelforce, though, let me know via Twitter that their software allows free form entry of addresses outside the UK, so the error lies with the sender; and I suppose it's not completely out of the question that somebody in the UK with an unusually high knowledge of European province names had typed the name in incorrectly.

To give them their due, Parcelforce answered all my tweets.When I suggested that they introduce basic validation into their address capture software, they suggested that the expense for validating every address in the world would be prohibitive.

It wouldn't even be possible. But you don't have to go the whole hog, from no validation to full validation. Too many businesses think that way. There is a lot that you can do to test an address which is very easy and very cheap and very effective.  You could, for example, test the basic validity of a postal code - length, allowed digits and characters, format.  You could only allow the entry of a province which is in the country of the addressee.  All this information is easy to obtain online, and easy to program. Partial validation is easy, cheap and effective - no organisation should be scared of trying it.

Parcelforce's parting shot was that responsibility for the collection of the correct address lies with the sender, not with them.  Very likely. But Parcelforce is owned by Royal Mail, a national postal authority, who must understand the importance and value of address validation. Fobbing the blame off onto each small business using their service is a bit lame.

The package arrived (late, as I said, because GLS claimed not to have been able to find my address first attempt and didn't want to bother contacting me to help them - I suppose if you're in a van with eyes front and not wanting to slow down to check building numbers it's as good an excuse as any), but it did highlight again how simple actions to verify basic address elements can be in everybody's interest and not the great drag on resources too many people imagine it to be.

Friday, October 17, 2014

Is the IAIDQ dead?

A recent post by Daragh O Brien about the International Association for Information and Data Quality (IAIDQ) and its future got me thinking.  I’ve never been deeply involved in the IAIDQ, unlike Daragh, but I was also a charter member, and I have experienced a definite change recently. Or, perhaps, experienced that I was no longer experiencing anything, if you get my drift.

Many of the people I knew who were involved in the early years of the IAIDQ have retired or moved on, and requests to me for information, articles and so on from the current leadership have dropped to nothing. Which may not be surprising, as they probably don’t know me from Adam. Indeed, as Daragh points out, a new edition of the Journal has become as rare as a web form which can correctly collect address data from more than one country (i.e. almost non-existent!). Members used to have a vote on members of the committee – that seems to have been quietly dropped too. 

I know that Daragh won’t agree, but I began to be concerned when the organisation started to busy itself creating Information Quality Certified Professional” (IQCP) qualification. What’s it for? I am a firm believer in educating people about data quality, but I don’t see how a qualification is a useful part of that apart from filling up space on a CV. I have no idea what the qualification entails – my services when it was being formulated weren’t required – but my impression is that it deals essentially with theory and not practice.  And it’s clear to me that those at the doing end of this data quality thing have no better understanding of data quality and how to achieve it than they did ten years ago. In fact, as businesses perceive that they need to obtain and manage ever larger amounts of data, even though often they don’t, accuracy and quality are diminishing – gather enough data and take a swipe at it, and you’ll hit a few targets on the way. Maybe the IQCP qualification is useful for some, and shouldn’t be harmful for others, but it does seem to me to have become the central focus of an organisation that should be doing more than counting the number of paying students they can hustle through an exam.

I don’t know if the IAIDQ is dead. I’m not close enough to it and they seem not to want to be too close to me. But one thing I do know. Much earlier this year I received an e-mail requesting that I renew my membership. Instead of immediately doing so, I cogitated on what I was getting out of the IAIDQ (nothing I could think of) and so, in straitened times, I decided to put off the decision on whether to renew until they sent a reminder.

I’m still waiting.

If an organisation dedicated to data quality can’t manage its own data, doesn’t that say something?

Wednesday, September 17, 2014

Sod's law

A recent blog post from a data quality company, which shall remain nameless, was in the form of a quiz. Could the reader spot the errors in the address formats from various countries?  The idea was good – trying to get the reader to appreciate that these differ throughout the world, whereas most companies think they are the same for everybody.  Unfortunately Sod’s Law intervened – the “corrected” addresses were mostly not, and the post was full of inconsistencies.  A quiet tweet in their direction and the post was removed. 
And that was the end of that. 

Except that this is a very common topic for blogs from data quality experts and providers alike. “Look”, they say, “this is wrong, and this is right, and we help you get from the state of being incorrect to the state of being correct.” 

Again, all good stuff. But there are a couple of important points which I rarely see addressed in those posts.

Firstly, correct according to whom? According to the local postal services? According to the local government? According to the emergency services? According to the bible of Graham Rhind? According to local cultural norms?  Although postal services are often the managers of street address files, they may not originate that data, and increasingly alternative resources, such as land registries, are becoming available to use instead. Often “correct” is taken to mean the form that an address takes in the local postal address file, if one exists. Those files are often held for a single purpose – to facilitate the efficient handling of mail – and, as postal authorities face the same problems of data quality and management as the rest of us, they may differ substantially from how the local populace actually write that addresses.  The data may, for example, be stored only in capital letters, without punctuation and without diacritical marks.

The fact that postal address files are used primarily for mail delivery brings me to the second point I miss when companies talk about what they are able to do – what is the address to be used for?  The blog post I mentioned suggested, for example, that an arrondissement (district) of Paris should be added to a French address and a county added to a British one.  We know that a county isn’t required in a UK address used for mailing, provided the postal code is there, and an arrondissement is not a requirement in a French address on a letter, especially as that information is also included in the postal code.  But if that address is being used in a travel guide, or on a website to show a business’s location, or to provide a route description for a person, then the additional data improves the usefulness of the address information and won’t be wrong in the address unless the different pieces of information don’t match (for example, if the wrong county information is provided).


I look forward to blog posts and articles about address data glitches. But is it time to move on from postal address files being regarded as the (only) holders of the golden record?

Thursday, July 10, 2014

Do I want online advertising to be relevant to me?

I recently read an article in Database Marketing Magazine by Paul Kennedy about the myths and reality of data (online version here). In it Kennedy suggests, and I paraphrase, that consumers would rather see offers and advertising online which is of relevance to them than generic advertisements, a point often made. Is this assertion true?

I don’t have any figures which support or refute this, but naturally the answer to a question depends on the question being asked. I suspect that given a choice most people would simply rather see less or no advertising than relevant advertising, or would rather see advertising of any type which is easier to distinguish from content than what is currently on offer. But most people also understand that the current financial model for online content is to provide it for “free”, paid for by advertising and often in exchange for people’s personal data. Without advertising the larger online companies wouldn’t be so rich and those of us with a smaller online presence wouldn’t still be in business.

Regardless, I’m not one of those who wants to see relevant advertising.  And I’ll tell you for why.

When I receive mail, or an e-mail, from a company, then I like the offer to be relevant to me, to be of interest, because I am offended by the waste involved, in time and resources, when it isn’t. But when it isn’t I can easily take action. I can dispose of the communication, which is a separate unit which I can choose to pick up and read when I want to, or discard, and then forget about. In many countries legislation exists which would allow me to turn these communications off. When the advertising block starts on the TV, I can turn it off, turn the sound down, or walk away for the duration. The advertising is isolated from the content (though increasingly less so), and that gives me, the consumer, the power of choice.

Upselling in mailings, such as with orders or statements, has been around for a while, but at least it is generally in distinct units – I can discard the guff and concentrate on the content. Up to now no company has tried to upsell to me on the same piece of paper as the invoice etc. with which it was enclosed, and let’s hope that that doesn’t happen.

Online advertising is different.  It is pervasive and invasive. It doesn’t form a separate unit which I can view or ignore, as appropriate.  It is woven into any content that I have actively sought out, and it is becoming increasingly difficult to distinguish as advertising. It is intrusive and sometimes so invasive that its purpose is defeated. A well-disguised audio advertisement on a page will have me backing out of that page as fast as my mouse can reach the button, surely to the detriment of the content provider. No legislation exists to allow me to view my content without advertising. It’s the equivalent of being sent a bank statement and then trying to find and view my account balance amongst the advertisements for fast cars and Ukrainian mail-order brides.  Unthinkable offline, but run of the mill online.

Online advertising is often dishonest. It lies or disguises itself as content to attract my click which, whilst profitable in the short-term for the pay-per-click provider, won’t help a brand in any way in the consumers’ eyes. I’ve seen pop-up advertisements in mobile apps with either no close button or one which is so small that a human finger will often miss it.

Emphasis on what adverts are shown is placed on the person viewing a page, which is why online advertisers are so keen to find out all they can about you and I. Why there isn’t more emphasis on the content we are looking for and looking at is a mystery to me.  If I’m looking at a page of reviews for hotels in London, then advertisements for hotels in London would probably be a better bet to get my click than ones trying to sell me a lawnmower. Once I leave those hotel pages and move on, though, I don’t want to be continuously subjected to adverts for hotels in London – that was then. I’ve moved on. Shouldn’t the advertising move on with me?

When I go online to look for something, a new watch for example, then I would like to see information about watches when I’m looking for it. Just as I would choose to go to a jewellers to find a watch when visiting my nearest shopping centre. Once I’ve left that shop/search, though, do I still want to be constantly marketed to about watches? Do I want to read about watches when I’m shopping for a fire extinguisher, or reading the news, or chatting to friends? Why would I welcome that distraction? Fine to see something while I’m looking for that product – it’s fair game that, if I’m looking for a watch you want me to buy yours – but afterwards? There are tracking cookies, more like stalking cookies actually, which keep presenting the items you viewed in one site on other pages you might visit.  Amazon does this.  It’s like walking out of the jewellers and having somebody follow you shouting a constant refrain of “BUY THE WATCH!  BUY THE WATCH! YOU KNOW YOU WANT TO! BUY THE WATCH” until you either give in or, like me, find and change the tracking preferences for that retailer.

So, as advertising is there and isn’t going away, do I want the advertising I see online to be relevant and “interesting” to me, in the same way as with direct mail?

No.


I’m clearly not the target of most online advertising, which is aimed at people who are as lax with their purse strings as they are with their personal data, but I don’t want online advertising to be relevant to me because, if I can’t choose whether and when to view it, then I’d like to be able to block it out as easily as possible. Whilst the pages I view are full of advertisements for cars, singles matching sites, holidays in the sun, football tat and flat rentals, in language(s) I don’t speak and none of which have any relevance to me at all, I can concentrate on the site’s content without distraction. This also reassures me that either companies haven’t got much personal data about me, or they don’t know how to use it. Either way, that’s fine by me!

You may have it. But do you know it?

A country may have a postal code, but is it used? Do people know there is a code system, and, if they do, what their code is? And when is it safe to make "postal code" a required field for forms for that country? Read more here.

Monday, June 23, 2014

Worts' Causeway

The potential for data to be corrupted and polluted increases as  it gets passed through interfaces and contact points, and as it passes from process to process and from system to system. This makes data hard to keep clean/ Those of us in the data quality world often hammer at the point that getting the data right at source is the ideal for a high level of data quality. But not all data is correct or standardised at source. When the originator of that data, as much prone to data quality defects as the rest of us, can't decide on what form it takes, what chance for getting it right in your systems?

Read more in my blog post here.



Sorry, you are an invalid character


In a bid to promote equality, and bringing it more into line with other European countries, a proposed new law in Belgium would automatically assign a baby the surnames of both the father and the mother (in that order) instead of only the surname of the father. However, the parents may also choose to give the child the surname of either of the parents, of to have the mother's surname precede that of the father.

Read more in my blog post here.