It's not hard to remember a different, strong password for every website

Let's take it as a given that it's a good idea to have a long password with upper case, lower case, numerals and special characters. Let's take it as a given that it's a good idea to have a different password for every website, and the main reason people don't is because it's very difficult to keep track of them all, and too much mental effort every time you need to sign in.

Your choices are:

  1. Use the same password for every site and hope nobody hacks it, and then uses it on all your other websites.
  2. Use a password wallet service and hope they never get hacked (NOT a given!), or nobody finds out the one password you use to sign in to it. 
  3. Find a way to have a different password for every site.

I choose #3. You don't need to memorize 500 passwords; you need to memorize one set of rules that allows you to easily mentally calculate your password each time. Here is one example; I use one just like it, except totally different.

1. Memorize a list corresponding to letters of the alphabet

This may seem daunting, but it's surprisingly easy. Within a week, you're able to recall them instantly with no problem; it would be hard to remember 26 random words, but alphabetizing them fools the brain into giving them structure, and structure is easy to memorize.

Our example:
aardvark, bear, camel, duck, elephant, fox, giraffe, hamster, etc...

2. Transform them so they are not complete words

xkcd notwithstanding, it's not a good idea to use complete words, because one hacking strategy is dictionary-based. There are many ways you could transform them, swap out some letters for others, remove all vowels, truncate them to the second vowel; in our case, we'll just take the first three letters so it's easy to follow.

aar, bea, cam, duc, ele, fox, gir, ham...

3. Replace letters in the target website

Use a non-obvious pattern. In this case, we'll take the first four letters of the website, but in reverse order. Our example website will be cabernet.com (it doens't exist... yet), so the letters are e-b-a-c and our code is now:

elebeaaarcam

4. Add some rules for capitalization, numbers and special characters.

The sky's the limit here, we already have a pretty good password, so you can limit the complexity of these rules so they're easy to implement quickly. For our example, we'll:

  1. capitalize the first and last consonant
  2. right in the middle, add 858 if the website ends in .com, 636 for any other TLD (I just took the easily remembered 747 and shifted it up or down a digit)
  3. at the end of the word, add %$# (that's the special characters above 543) if the website name begins with a vowel, #$% (the same, reversed) if it begins with a consonant.
So our example is now:

eLebea858arcaM#$%

This scores 100% on passwordstrength.com, and most importantly, if a hacker finds it out due to a vulnerability on a website you're signed up for, and over which you have no control, they don't now have all your passwords or an easy way of figuring them out.

This is a bit of mental effort, traded for a lot of security. And it's a lot less mental effort than it seems; the human brain is really good at remembering and implementing repetitive rules. My algorithm is somewhat more complicated than this example, and I never have to hesitate more than a second or two, and I always gets it right.


The five commandments (and fifteen footnotes) of data visualization

The five[1]  commandments[2]  of data visualization[3] 

I'm nobody special in the world of data viz; I have no profound observations or innovations to add to those of the likes of Edward Tufte, Hans Rosling, Hadley Wickham or Mike Bostock; but I think I have a little common sense and boots-on-the-ground experience when it comes to the more mundane, journeyman work of making a PowerPoint slide and being proud it doesn't use Comic Sans[4] . (I use Python now, I'm never going back.) By all means, if you have the time and luxury, concern yourself with data-to-ink ratio[5] ; but before that point, here are a few tips to help ensure there's any ink at all.[6] 

(Comments, suggestions, dissenting opinions and, especially, corrections are very welcome, don't be shy. Yes, I'm saying don't be shy to the Internet. Stand back.)

1. A graph is like a paragraph.

A visualization should have more to say than a sentence or a sparkline, but less to say than a short story or, well, the raw data itself. A data visualization needs to strike a balance between saying too much and saying too little.[7]  A good rule of thumb I use is: if one really well-written footnote might help someone understand the graph a bit better, I've done my job right, and often I don't even end up using that footnote. If the footnote is absolutely necessary or two footnotes present themselves, or if not even a five-year-old would see any value in a footnote whatsoever, then maybe the scope of the graph is wrong.

2. Visualization is translation

The Italians have a saying: "Traddutore, traductore", roughly "Translation is treason". Creating a visualization is translating raw data into another medium, and it involves loss of information, and it involves choices, hard choices, desperate choices [8] . Think of it as describing a movie to your significant other. ("I know you don't like action movies, but it was so cool when Schwarzenegger threw this grenade..."[9] ) You can't describe the entire movie, nor does your audience want you to. They trust you to make the choices necessary to get the pattern hidden in the data across with skill, clarity and integrity. Which leads me to...

3. Visualize with integrity

When you were a child, someone must have told you honesty is the best policy[10]  , and they were right. When in doubt, act with integrity. Actually, when not in doubt too. Actually, especially when not in doubt: if you're not wondering if you're making the right choices when you're deciding how to visually present your data, you're doing it wrong. [11]  Basically everything I need to say about integrity, Fox News has said more eloquently (and just slightly less intentionally) than me[12] :


4. The second-worst outcome is for someone not to understand your graph.

There are lots of very complicated visualizations out there. They certainly have their place, but they probably should be tackled by the elite among us.[13]  But even with more modest goals in a visualization, it's quite possible for the message to get muddled in the medium [14] . Boxplots are brilliant. I absolutely love boxplots. But so few people understand them, more time gets spent explaining how they work than on trying to understand the actual data. Of course, it all depends on your audience, and your mileage may vary.

5. The worst outcome is for someone to MISunderstand your graph.

What's obvious to you as the creator might not be so to your audience. Always ask yourself what a fresh pair of eyes might understand from your graph. If necessary, go and find a fresh pair of eyes. [15]  When someone misunderstands what your visualization presents, no matter how obvious you think it is, you failed. ("You had one job!") You are intimately familiar with your data: others are not, and first erroneous impressions are hard to erase, and people get upset when they need to rearrange their brains. Don't scoff, you do it too.

6. Don't limit yourself to your first idea.

Like "This will be a list with five items on it."[16] 




Footnotes:

[1] I just read an article that claimed the most successful web content is list-related, has an odd number of items in the list (exception: 10 items), turn browsers into subscribers and has a (relevant) photo of Jennifer Lawrence. I've got at least three of those criteria covered. Return to article.


[2] I really don't have any basis to be commanding anyone, but "suggestments" just didn't have the same ring. Return to article.

[3] That's what you're supposed to say now, instead of "graphing". For one thing, graphs (which I used to draw on graph paper with a ruler and a bendy curve thing) are now called charts; graphs look like this: (O brave new world that has such network diagrams in't.) Return to article.

[4] Go for Papyrus or Mistral instead. Return to article.

[5] But if you want a lesson to keep in mind from the gurus, you could do a lot worse than Edward Tufte's idea of data-to-ink ratio. In most human endeavours, less is more: art, design, psychology, John Travolta movies, you name it. Return to article.

[6] Or, in most cases, pixels or toner. Don't be pedantic, you know what I mean. Return to article.

[7] Here's where a more trite writer would whip out pairs of words that are oh-so-similar-yet-oh-my-god-crucially-different: Simple, not simplistic. Complex, not complicated. Focused, not targeted. Zuul, not Dana. Return to article.

[8] I would firmly like to state, so that there is no misunderstanding in the future, that I fully support a data visualizer's right to choose. (And they said I could never work an abortion joke into a data viz how-to list. Who's laughing now? Um, probably nobody, it was a really bad joke.) Return to article.

[9] Many times I sat through my girlfriend giving me these rundowns, yet she wouldn't even listen to me gush about the Zooey Deschanel movie I'd just seen in the other cinema. Return to article.

[10] Ironically, as adults, we tend to use the word "policy" mostly to describe the work of politicians, insurance agents and businesspeople. I'm not saying they're not honest, but... actually, yeah, that's pretty much what I'm saying. Then again... I might be lying. Return to article.

[11] If you're agonizing about the choices so much that you curl up into a foetus of self-doubt and vow never to make so much as a bar chart ever again... well, you're doing it wrong then, too. Return to article.

[12] I have seen many equally egregious though perhaps less deliberate non-zero graph origins over the years: don't do it if you don't have a good reason to, and if you have an excellent reason to, don't do it just the same. Return to article.

[13] Preferably when very, very drunk. I'm looking at you, Nathan Yau. (I don't know anything about Nathan Yau or his personal habits, but God, that was fun to write. I'll apologize for my nasty whimsy by suggesting you buy his books.) Return to article.

[14] McLuhan might've moaned at my misuse of a memorable maxim, mmm? Return to article.

[15] Hopefully still attached to a head, they tend to be more useful that way. Return to article.

[16] That goes double for footnotes. Return to article.

How not to make an infographic

I became unexpectedly unemployed yesterday, and since I don't believe in long mourning periods (or poverty) I started my job search right away, and came across this infographic. Let's be fair: there are far worse infographics out there. But my version of human nature somehow gets more perturbed by almost-competence than by abject failure; I suppose, knowing nothing about the creator, that in my head I'm blaming them for not trying hard enough. Well, if the creator happens to come across this, I totally don't want to hurt your feelings (much), you just need a little more practice, as do we all. 


The Meaning of The Game of Life


When it was family board game night in the '70s, I always picked The Game of Life or Careers (or sometimes Masterpiece ... I was a weird kid), not Uno or Aggravation. I guess I liked the complexity; I even found Monopoly too predictable (it's a Markov chain, after all). My love of complexity was given a mainline overdose in 1979 when my dad became the first person in our town to own a home computer, a Texas Instruments TI-99/4 with a whopping 4K of RAM. (That's 1/16 as much as a Commodore 64.) He also had a subscription to a programming magazine, so every month I would read through mostly incomprehensible (to me) lines of badly-typset BASIC (a different dialect than TI BASIC, which caused me no end of grief), and occasionally laboriously type some in. There was Eliza the automated psychotherapist, the Lissajous fractal generator ... and Conway's Game of Life.

In 1970, British mathematician John Conway invented a set of rules for cellular automata that would allow you to start off with any combination of black and white squares on a grid like a chessboard and by applying three simple rules determine whether in the next turn a square would stay the same or change colour. Since computer time was too precious to be wasted on mere experimentation, Conway designed Life with pen and paper and only computerized it when he knew he would get a passable result. And what a result! While most cellular automata (a concept invented as a side project by scientists working on nuclear bombs) until then had been uninteresting, quickly gravitating to all black or all white squares, Conway's developed long, sometimes extremely stable cycles, with some shapes persisting, some cycling, some even spawning other shapes, looking like birds or stars or spaceships or even cells... it was a perfect storm of mathematics, armchair philosophy and whimsy at a time when computing was just starting to be able to accommodate all three.

There are many websites (and a humungous wiki with hundreds of named patterns) devoted to The Game of Life, but it seems to have lost its place in the geek imagination over the last ten years or so: most of these sites use Java applets that are routinely disallowed by modern browsers. To see Life in action, you pretty much are now stuck with Animated GIFs or downloadable freeware. Life is the first computer program I actually understood. I typed it into my poor overworked computer (it got so hot I actually burned my elbow once), and when it ran, the 7 or 8 second pause for recalculation between iterations didn't bother me. (At the time, everyone compared computer pauses to how long it would take a human to do with a slide rule and graph paper, so it was miraculous.) Then in the next issue, there was a revised code of the same program that ran over twice as fast by holding only three lines in memory at a time instead of the entire array. My mind was blown and I sat down with a pencil for hours and marked up the code to understand why this was so.

Then I gave up coding for 30 years to study opera and work in underground journalism, but that's another story.

(I've come full circle with keyboards, however: I now use the first chicklet keyboard I've had since 1979. Oh, and in the meantime Steven Wolfram claims to have reinvented science (or something, I only understand about 5% of it) using cellular automata with catchy names like Rule 90, not to be confused with Rule 34... don't worry, the link is safe for work).

To celebrate The Game of Life, I've rewritten some BASIC code using Joshua Bell's terrific javascript emulator (he was kind enough to help me implement it, since my BASIC is way better than my javascript). Enjoy, and marvel at what passed for computer graphics when Jimmy Carter was president. If you're a programmer, feel free to look at the code and smirk at the LET syntax, lack of ELSE statements and highly vulnerable GOTOs (I wrote so many infinite loops it was practically my breakfast cereal ... unlike Mikey, who preferred Life).

Note: Blogger seems a little finicky (and/or I'm a little incompetent) when it comes to displaying javascript. If you don't see a big black box above this sentence, click here or here. But not here or here. And definitely don't click here.

Use of the f-word in Eddie Murphy: Delirious

I had planned to take a break from blogging during the holidays, but today I saw this post on reddit about the use of the f-word in movies in the dataisbeautiful subreddit, and I was inspired. The top movie on the list I had seen was Eddie Murphy: Delirious; I was 13 when it came out, but nobody I knew had HBO, so my best friend and I had to wait till it showed up in the Betamax tape rental place. We made a lo-fi audio recording (a microphone held up to the TV speaker), and soon had it memorized and spent several years quoting it in all sorts of inappropriate situations.

So, let's break down the use of the f-word (I admit, I'm being a total wuss, Google hosts this blog and I'd rather not deal with any automated fallout from using profanity, so I'm going to asterisk out all the naughty words) during the movie. Some simple poor man's calculus (for each use of the word at time x, y equals the inverse of the average of the times of the previous and next use) shows the clustering of swearing during different parts of the film:


It would be great to know what parts of the movie those clusters correspond to: if you go to the bottom of the post, there's a reversed version of the graph that allows you to see the dialogue (lightly Bowdlerized, again, I'm sorry) line by line.

I've been learning how to do Natural Language Programming in Python, and while I didn't bring out the big guns, I thought it would be interesting to look at some of the simple patterns in word use in the movie: 


Normally I would use a stop list to remove common words like "the" and "and", and a corpus to compare word frequencies, but I think the raw data is the most informative perspective, showing how the profanity rivals the most common syntactic words in Delirious. Here are the top N-grams (words that appear side-by-side):



I'm a contributor to the FullMovieGifs subreddit, so I couldn't resist the temptation to make one of Delirious. Hopefully Google doesn't OCR these things; if you want to see it larger, click on it.




Finally, here's a big, vertical version of the first graph in the blog, which you can mouseover to read the lines of dialogue (is it still called dialogue when only one person's talking?) to your heart's content. If you can't see a really huge graph right underneath this sentence, click here to see it.


I think I'll be hearing from my mom about this post.
Update Jan. 1, 2014: Whaddaya know, my mom was fine with it.

Weight of small change in USA, EU, UK and Canada

I'm interested in how much the SMALL change weighs, I don't want to get into the dollar bill/coin debate.
The graph is interactive, feel free to click and hover. [Blogger seems to be finicky with javascript; if you don't see a big interactive graph right underneath this sentence, click here.]
Bottom line: Brits need good pockets.


I started learning javascript a couple of months ago, and I'm comfortable enough to be able to lean heavily on a package and wrangle the API to give me what I want. Today Highcharts, tomorrow, D3.js!Coin weights are taken from Wikipedia.
If anyone prefers to see a simple non-interactive image, click on this:





Boston Celtics retired jerseys by year: when will they run out of numbers?


I'll admit I'm not a huge sports fan, but I am a huge numbers fan, and sports produces a lot of those. It also produces a lot of analysts: after all, there's lots of money riding on much of these numbers. So it's a bit of a challenge to find something original, and by definition it's going to be a bit frivolous.

It occurred to me that if teams keep retiring numbers and don't expand the pool of possible numbers, eventually they will run out. A bit of Googling revealed that the Boston Celtics have the most retired numbers of any major professional sports team. The NBA allows 100 numbers, from 1 to 99 and 00; they've retired 21 in the past 40 years, so a simple linear fit shows that at this rate they will run out in a couple of centuries.

I wouldn't worry about this problem too much; the Celtics have already shown how to solve it. When they retired Jim Loscutoff's jersey, he requested that they not retire his number (18), so their banner reads "LOSCY" instead. Later, Dave Cowens spoiled the gesture by wearing the same number and having it retired.

It occurs to me that I've seen these kinds of stepwise and extrapolation graphs on xkcd (e.g. here and here), except of course Randall Munroe is much better at them than me. So I decided to do a little tribute and rework the first graph xkcd-style using Dan Foreman-Mackey's xkcd D3.js template. My javascript skills being what they are, this was by far the longest part of this project; but it was a labour of love. I hope everyone will forgive me.


Population of Canada by latitude



Update: here's my final edit of the chart; I think the city labels are much less misleading now. I've come across a much more fine-grained data set, albeit from 1995; you can see it in my Nov. 27, 2013 blog post.



Here's the original, which seemed to imply that the bars were only made up of population from the indicated cities, whereas the bars indicate the population of the entire country at the same latitude of those cities:



A co-worker and friend happened to mention that Vancouver was further north than Montreal; I sort of knew that, but I was surprised to find out it was 400 km further north. So I was curious, and tried to find a histogram of Canadian population by latitude; maybe my Google fu was lacking, but I couldn't find one, so I decided to make one myself.

Little did I know what I would discover; that data is not easy to obtain. There is lots of population data available for download from the Statistics Canada website, but it does not contain geographical coordinates, and StatsCan uses its own defined areas called census subdivisions. They have available for download geographical boundary files, but they would have required an amount of computation rather disproportionate to the task of simply determining latitudes.

Luckily, StatsCan also makes the population available by Forward Sortation Area, the first three letters of the Canadian six letter postal code, e.g. the FSA of the Canadian parliament at postal code K1A 0A9 is K1A. So now it was just a matter of finding out the latitudes of FSAs or postal codes. Simple, right?

Wrong. Canada Post considers its postal codes intellectual property subject to copyright; a license to use and analyze it costs $892 a year for StatsCan's info, and over $5000 for many business products. They are suing a website for providing information on postal code geography. Universities used to be able to access Canada Post's geographical data, but no longer. I work for a university, and the reference library has someone who is able to take the publicly available ArcGIS files and determine the centroids using the expensive proprietary commercial software for which the university has a license.

So: the population data is divided into 1600 FSAs, which is pretty decent resolution. The centroid (geographical center) for most postal codes fits reasonably well within the 0.5 degree latitude (about 55 km) resolution of the graph, except of course for the very large FSAs the farther north you go. But in any case, these areas would have had to be aggregated somehow to even be visible on the scale (for example, if if the northernmost FSA, X0A, were spread out among its 14 degrees of latitude), so I think this is a reasonable compromise.

A note on the city labels: I tried to give the largest municipalities that contributed to the population in each bar of the histogram as an aid to understanding, not as a systematic data set. This became difficult for some of the larger FSA's; it was difficult to match the latitude of a town with the latitude of the centroid of its FSA. So in some cases, I may have used a town with a population of 2,000 when there was a town with 3,000 people at the extreme north or south of the FSA. And a note about Edmonton: it straddles two bars because the center of the city is almost exactly on the demarcation, 53.5 degrees north. Edmonton is a bit smaller than Calgary, but there are other sources of population in each latitude than the city mentioned, so do not draw the wrong conclusion from the size of the bars.

You can peruse the data I used in this Google Doc.

Comments are welcome, even, nay especially, critical ones.

EDIT 2013-10-16 14:49 GMT: Montreal straddles the 45.5 degree latitude, and by marking the 45.5-46.0 bar as "Laval", the graph appeared to be indicating that Laval had a larger population than Montreal. I've explained how the labels are generated, but it's an obvious conclusion to draw from a glance at the map without reading the methodology (and the methodology had to be tweaked for Edmonton and Montreal, which straddle the cusps of the graphs, and the centroids of the FSAs are problematic to begin with). Clarity is the most important thing, so I've updated the bar to read "Laval & Montréal". Thank you to the commenters in Reddit's dataisbeautiful forum for pointing this out.

EDIT 2013-10-16 15:33 GMT: When you're wrong, you're wrong, and I was wrong. My labels were utterly misleading. Now I have put the major contributor AND every Canadian city with over 100,000 population on the graph. I had intended the labels just as a geographical reference, but I definitely did not think through what fresh eyes coming to the graph would think.

EDIT 2013-10-16 21:53 GMT: These labels are really getting me in trouble. I produced the graph first without them, but I envisaged a torrent of "You should have indicated where these people live!" I've removed the most northerly ones, because again, they're misleading. Lesson learned: less is more.

EDIT 2013-10-16 22:41 GMT: Added hi-res version without labels. I think that's enough editing today. Enjoy! And thanks for all the feedback! The vast majority of it was very constructive, it's appreciated.