DDJ - Resources http://datadrivenjournalism.net/resources DDJ - Resources en support@ejc.net Copyright 2015 2015-07-27T08:13:45+00:00 Use flowcharts to avoid getting stuck in data journalism projects http://datadrivenjournalism.net/resources/use_flowcharts http://datadrivenjournalism.net/resources/use_flowcharts#When:08:13:45Z All journalists are quite comfortable coping with the problems, big and small, of a written story. You start with an initial idea, then do your research, gather information, put it into a form fit for publication. 

The same process in a data story can pose quite a number of obstacles. Often, people get stuck. So, what do you do when the data is hard to come by? How do you avoid to get stuck with an overly complex set of data? What is the best visualization to choose later?

Planning ahead can help to avoid that. Similar to writing a story or creating a video work on a data story is facilitated when you have a clear step-by-step plan at hand. 

This is why flowcharts can help a lot finding your own path. Recently Journalism.co.uk published story linking to a blog post to project flowcharts which can be a big help - for your own work and in order to communicate what you intend to do to others in your newsroom (including superiors).

One example of such a flowchart comes from science journalist James Gaines. He writes: „I seem to constantly get stuck. Data journalism is really useful and powerful, but it does require a bit of technical know-how to perform. More than any other form of journalism there definitely seems to be a right way to do it and a wrong way to do it. So in order to help myself when I’ve gotten stuck (and to help me remember what program to use when) I’ve made myself a flowchart.“

See the flow chart he created for his own projects here.

(The best advice in this plan is to talk to an expert early in the project in order to avoid drawing the wrong conclusions as to the quality of the data. Pick up the phone early in your project and reach out!).

Later, as a response, Paul Bradshaw picked up on the topic and created his version of a flowchart, particularly for the process to gather data.

The key point here is that planning makes it much easier to go through the steps of your particular project: Almost all data projects go through phases - from an idea or dataset to gathering the numbers, to cleaning, to visualizing and then on to publishing. No two projects are exactly the same - so creating your very own flowchart might be extremely helpful to chart your own course.

Equipped with such a plan, the range of tools you might use is easier to put to work. There will be turning points where you just know that at a point the best option will be use Excel, Open RefineR, Kimono or Import.io for the data gathering and cleaning. But having a visualized plan will help a lot to make decisions at turning points and stages of your particular project.

]]>
2015-07-27T08:13:45+00:00
Training: Data Visualization and Infographics with D3 http://datadrivenjournalism.net/resources/data_visualization_d3 http://datadrivenjournalism.net/resources/data_visualization_d3#When:07:20:25Z The Knight Center for Journalism in the Americas is preparing for another D3 online course starting on August 17. As of late July 2015 registration is still open.

D3.js, created by Mike Bostock, is currently one of the most popular and extensive data visualization options. The one thing keeping many journalists from using the library in daily work is the relatively steep learning curve - to use D3.js one firstly needs to learn about how to combine data and visualizations for publication.

The online learning opportunity will be hosted by Alberto Cairo and Scott Murray (Author of „Interactive Data Visualizations for the Web. An Introduction to Designing with D3", O’Reilly Media, 2013).

The syllabus of this new course covers principles of infographics and data visualization in six modules.

Note: The registration fee for participation is $95.
See link to register below.

Course overview:
http://journalismcourses.org/D30815.html

Syllabus:
http://journalismcourses.org/D30815syllabus.html

Registration:
http://journalismcourses.org/login/index.phpexp

]]>
2015-07-26T07:20:25+00:00
Speaking in tongues http://datadrivenjournalism.net/resources/speaking_in_tongues http://datadrivenjournalism.net/resources/speaking_in_tongues#When:06:31:25Z In Europe a constant flow of information across borders means that texts are translated back and forth. Surely translations this has an effect, however subtle. Why are some topics gaining in publicity and others don’t? What effect does the translation process of news, facts and data on how topics are treated in national contexts?

Who is doing the translation? How is the process of transformation being checked? How do we track the subtle differences in meaning and understanding? Given the many languages spoken in EU member countries translations play a so far mainly unmeasured role on how and when topics are noted, understood and acted upon. Jessica Mariani, a PhD candidate from the University of Verona researched the topic for her PhD thesis.

Why did you choose this particular topic for your thesis?

Mariani: „While training as a Press Officer at the European Parliament in Brussels, I realized the importance of an accurate information flow across the EU and its geographic, linguistic and cultural boundaries.

Press Officers and journalists in the European context frequently translate news texts from English/French into their own language and vice versa, although they don’t often qualify themselves as professional translators. Semiologist and writer Umberto Eco claimed that “Translation is the language of Europe”; communicating EU activities and aims has become a challenging task for Press Services and European media professionals in such a multilingual context.

At present, News Translation is an ordinary task for press officers and journalists, but it still remains under-investigated by academics, with reference to translation processes and cross-linguistic transfer. Thus, research and ethnographic investigation of journalism everyday practices might contribute to outline the role translation has gained in reporting the news and measure the risk of misinformation across languages and cultures.“

How did you investigate differences, what main theories were investigated?

Mariani: "In order to support a thesis, one needs to firstly base the research on case studies and provide evidence. In this case, I have chosen McNelly’s (1959) Theory of News Flow as my starting point. The news flow goes through three phases, from the institution to the press service, from the press service to the media and finally, from the media to the readership; this means that a news-text gets re-translated several times. Furthermore, Translation Studies as a discipline provides useful tools for linguistic analysis; Critical Discourse Analysis, whose forerunner is scholar Fairclough, enables researchers to analyze translated texts not only from a linguistic perspective but also with reference to the “actors” involved and the context itself. Translation is not only based on linguistic equivalence; as stated by Edward’s Sapir (1956) theory about linguistic difference: “No two languages are ever sufficiently similar to be considered as representing the same social reality. The worlds in which different societies live are distinct worlds, not merely the same world with different labels attached”.

„The worlds in which different societies live are distinct worlds, not merely the same world with different labels attached”
- Source: Edward Sapir

Are journalists aware of the role translations have for reporting?

Mariani:  „What has emerged from academic research so far is that journalists do not usually define themselves as translators; instead, they prefer to be referred to as “multilingual journalists” or “international journalists”. But translation is not only based on literal linguistic equivalence and often involves cultural patterns that professional translators are usually trained to evaluate.“

What is your main suggestion to newsrooms?

Mariani: „Newsrooms should be aware that “translating” is a cultural process and if attention is not paid to certain cross- cultural elements, it can easily lead to misunderstandings and misinformation. What I will try to discuss in my thesis is also the eventual birth of a new professional role, the “journalinguist”, a media professional who possesses consolidated linguistic skills and news-sense.“

 

Brief bio
Jessica Mariani is a PhD Candidate in Media and Translation Studies at the University of Verona and her three-year research project entitles: “News Translation in European Context: Building a European Perspective”. After graduating in Journalism and Public Relations, she worked as an entertainment reporter at “Hotpress Magazine” in Dublin and Verona. She has recently trained as a Press Officers at the European Parliament in Brussels. Her interests range from European Affairs to Music and Entertainment, to Investigative and Data Journalism to Transparency and Civil Rights. At present, she researches the language of translated news with reference to the EU and investigates the role of “news translators” in newsrooms.

]]>
2015-05-09T06:31:25+00:00
Up to 100.000 Euros for projects connecting science and data-driven journalism http://datadrivenjournalism.net/resources/up_to_100.000_euros_for_projects_connecting_science_and_data_driven_journal http://datadrivenjournalism.net/resources/up_to_100.000_euros_for_projects_connecting_science_and_data_driven_journal#When:23:00:14Z The foundation recently published a call for proposals, the deadline for applications is June 15, 2015. Grants available can be up to 100.000 Euros.

Quote: "The idea is to initiate joint research and reporting projects which enable both sides to learn from each other and to generate new impulses for their respective activities.“

Teams who want to apply should have „at least one person from the realm of science/scholarship and one from journalism.“

Proposals will be evaluated by an international panel of comprising experts in the fields of scholarship and (data-driven) journalism.

Link:

Volkswagen Stiftung: Science and Data-Driven Journalism

 

]]>
2015-03-25T23:00:14+00:00
The Data Visualization Catalogue http://datadrivenjournalism.net/resources/the_data_visualization_catalogue http://datadrivenjournalism.net/resources/the_data_visualization_catalogue#When:07:00:39Z Why are bar charts so effective? When did the first bar chart appear? What do you need to look out for when using a bar chart in an article?

If such questions are of interest to you, head over to a cool project by Severino Ribecca, who is currently working on a series and collection of information. Severino is a graphic and information designer from the UK.

The Guide is useful for beginners and handy for more seasoned practitioners who need to sent a link to colleagues to make them understand the basics of charting better (which is a surprisingly frequent task).

On top of the catalogue Severino is currently publishing articles on single types of visualizations on Visual Loop. Here is a link to an article about Bar Graphs. A little digging will unearth similar articles on the site. They are concise, short, well illustrated.

Link:
http://www.datavizcatalogue.com

]]>
2015-03-15T07:00:39+00:00
The Structure of Data Videos http://datadrivenjournalism.net/resources/the_structure_of_data_videos http://datadrivenjournalism.net/resources/the_structure_of_data_videos#When:22:06:37Z We know video. We know infographics and data visualization. Of course, both can be combined. But, as always the know-how about the options, possbilities, the weak and the strong ways to tell a story makes a difference. The better you understand the form and it's variations the more likel is compelling output.

As a genre telling stories through charts isn't new. But it is an evolving field. Search a bit on YouTube and you'll find data videos for all kinds of purposes, which are centered around data. A paper, titled "Understanding Data Videos: Looking at Narrative Visualization through the Cinematography Lens" is worth being printed and studied. 

The authors conducted a study of data videos and the different elements used. In a second step they list up key elements of such narratives. Finally, the paper reports about an experiment, where the same data was used for entirely different story outcomes. Interesting and helpful if you want to create such videos yourself. 

Link:
Amini, Fereshteh et al: "Understanding Data Videos: Looking at Narrative Visualization through the Cinematography Lens"

]]>
2015-03-13T22:06:37+00:00
The Slicer: Cutting Pivot Table Data to Size http://datadrivenjournalism.net/resources/the_slicer_cutting_pivot_table_data_to_size http://datadrivenjournalism.net/resources/the_slicer_cutting_pivot_table_data_to_size#When:04:31:51Z By Abbott Katz, London-based Excel instructor and freelance writer, author of Excel 2010 Made Simple and spreadsheetjournalism.com.

The difference between a mistake and a shortcoming is in the packaging. If your math teacher tells you 2 plus 2 is 3, it's a mistake; if she tells you it's 4 but expounds her little equation in Sanskrit, it's a shortcoming. (And if you actually speak Sanskrit my analogy exhibits a shortcoming all of its own, doesn't it?)

And so it is with Excel's pivot table filter, an efficient and understated way to isolate and display a set of records from a larger whole that meet a specified criterion, while excluding from view the records that don't. Thus, for example, if you've been entrusted with a spreadsheet tracking campaign contributions by political parties (say in the UK), the filter will let you present only Labour or only Conservative contributions.
Of course, filters aren't confined to pivot tabling. You're all familiar with this button:

Screen_shot_2013-09-19_at_3.32.52_PM.png

Which when clicked, enables the user to extract a subset of records from its superset of data, e.g.

Screen_shot_2013-09-19_at_3.34.15_PM.png

Excel's pivot table Report Filter (called the Page field in an earlier life, a reference that suffered from shortcomings of its own) works similarly, if not quite identically, affording the user a means for pinpointing specific records from a position above the pivot table, as it were. Consider this unremarkable collection of data, which you can download here.

Screen_shot_2013-09-19_at_3.38.31_PM.png
 

Move the data through a pivot table and assign the fields thusly:

Report Filter: Name

Values: Test Score.

Clicking the Report Filter's down arrow unrolls the list of student names; the next click reports any selected student's score, e.g.:

Screen_shot_2013-09-19_at_3.39.49_PM.png

(Note the default mathematical operation, sum.)

That's surely straightforward, as it's meant to be (imagine 50,000 student test scores cascading down the spreadsheet, for example, instead of our diminutive ten, and the feature begins to make more sense).

That's all well and good, but in the interests of feature richness, Excel's 2007 release enlarged the filter's grasp, empowering it to return multiple entries. Click the Filter's Select Multiple Items option and proceed to tick the desired boxes e.g.:

Screen_shot_2013-09-19_at_3.40.46_PM.png

And you'll wind up with something like this, alas:

Screen_shot_2013-09-19_at_3.41.35_PM.png

The test scores - in the plural - are here indeed summed, but the (Multiple Items) entry in the Report Filter field is Excel's way of stating that you can't know exactly who those items, or students, are - and that's a… shortcoming.

And Microsoft knew it. So by the time release 2010 begged your attention a footnote of sorts attached itself to the Report Filter - the Slicer:

Screen_shot_2013-09-19_at_3.43.00_PM.png

The Slicer is an easy-to-use, curiously free-standing window on the contents of the Report Filter. In our case, if we drag Names away from its Report Filter berth and click PivotTable Tools > Options > Insert Slicer and tick Names, OK, in the Insert Slicers (note the plural reference; more on that later) dialog box, you'll conjure this tableau:

Screen_shot_2013-09-19_at_3.44.22_PM.png

Note the Slicer has inherited the previous Ed/Jane determinations, but in any case you've likely already understood the essential workings of the feature. Click any name, and the Slicer registers that student's score in the Values area; and you can earmark multiple selections by keying the ageless Ctrl-click tandem for non-adjacent names, or Shift-click to nominate contiguous ones (that is, Shift-clicking the first and last items in the span of names).

The Slicer's advantage, then, is in the first instance informational: parachute it upon the data and the user, and everyone else can glimpse precisely those items in the field designated for reporting. And because it isn't locked into place you can drag the Slicer about on the worksheet, and even copy and/or move it to a companion worksheet (just click a Slicer border in order to select it and apply either the good old Ctrl-C or Ctrl-X strategy), where it can continue to finesse the data (you can also delete a Slicer by selecting it by one of its boundaries and pressing Delete).

What the Slicer (and, for that matter, the standard Report Filter) won't do is enable the view below, one about which a student asked me just the other day:

Screen_shot_2013-09-19_at_3.45.23_PM.png

in which the names of the students feature squarely in the pivot table (it should be added incidentally that the Grand Total above is likely irrelevant; you're not likely to care about the total points a filtered set of classmates accumulate, but you might be concerned to determine their average, for example). If you want to realise the above view, you'll need to tow the Names field from the Report Filter into the Row Labels area. Once there, however, the Slicer can continue to sift the names.

On the other hand, one shouldn't get too carried away over that capability, because you can do much the same without the Slicer, by clicking the filter down arrow shadowing the Row Labels header:

Screen_shot_2013-09-19_at_3.46.15_PM.png

Now while it appears as if the Slicer was initiated into Excel's feature set in order to relieve the presentational problem we described earlier, the Slicer also offers up a couple of other substantive assets that could serve the user well. It's not a terribly well-practiced stratagem, but you can stack two fields worth of data into the Report Filter area. For example, download this large collection of school assessments conducted by OFSTED, the UK agency charged with inspecting the nation's educational establishments:

The workbook enables a good many interesting permutations, enabling the journalist to break out several types of assessment scores by type of school, school phase (i.e., primary or secondary), geographical positioning, parliamentary authority, and the like. For starters, then, we could dust off a pivot table bearing this initial structure:

Values:  Quality of Teaching (by Average)

Report Filter: Local Authority (education)

Type of Establishment

Now here's the problem. By clicking a particular Local Authority, say the London borough of Barnet, the accompanying filter field - Local Authority - does not automatically confine its data to those kinds of Authorities represented in Barnet. To illustrate: If I leave Barnet in place, and click the Foundation Special School establishment (look here for an explanation of this school type), the Values area yields:

Screen_shot_2013-09-19_at_3.49.53_PM.png

Nothing, and that's because there are no Foundation Special Schools in Barnet; and as such, one might have supposed the double-Report Filter above would have clued the user about that absence, or pre-empted the complication, but it doesn't.

But select the Local authority and Type of establishment Slicers instead and click Barnet in the former, and you'll see:

Screen_shot_2013-09-19_at_3.51.44_PM.png

That's more like it. We see that a selection in one Slicer has a conditioning effect on the other. The Foundation Special School establishment type is dimmed, reflecting that type's absence in Barnet. Click the Birmingham Local authority, by contrast, and Foundation Special School makes itself available. It's a cool and useful feature.

And there's one more generic filter tip you'll want to know about. If you return to the conventional Report Filter mode by dragging a field into its confines and see to it that the field shows (All) - that is, you haven't selected any particular field item - then click PivotTable Tools > Options > Options (in the PivotTable button group) Show Report Filter Pages…, and click OK in the Show Report Filter Pages window (which asks you to decide between filter fields, in the event you've thrown two or more of them into the Report Filter area). Once you've made your choices, Excel immediately turns out a new worksheet for each item in the filtered field - a series of mini-reports compiled around whatever fields you've tossed into the Values area. Thus, if I frame a Pivot Table with these constituents:

Report Filter: Government office region (make sure it shows (All))

Values: Quality of Teaching  (set, say to Summarize Values by Average)

and execute the Show Report Filter Pages routine, a parade of new spreadsheet tabs should expand the workbook, e.g.,

Screen_shot_2013-09-19_at_3.52.41_PM.png

with each - reporting the average Teaching score per region.

I don't know about you, but I've been so inspired by all this I'm having a slice for dinner tonight - of pizza, that is.

]]>
2013-09-20T04:31:51+00:00
Hate Spreadsheets Formulas? Meet Drag and Drop Data Analysis Tool ‘QueryTree’ http://datadrivenjournalism.net/resources/Hate_Spreadsheets_Formulas_Meet_Drag_and_Drop_Data_Tool_QueryTree http://datadrivenjournalism.net/resources/Hate_Spreadsheets_Formulas_Meet_Drag_and_Drop_Data_Tool_QueryTree#When:02:29:04Z By Daniel Thompson, software developer working on QueryTree and CEO of D4 Software.

During the course of this decade, the amount of digital information being stored by the human race will increase 50 fold. Every aspect of life now leaves some form of digital trail and this trend has implications for all areas of life, not least for those of us who seek to hold governments and large institutions to account. The job is changing; the stories are now hidden in increasingly large sets of data and to find them we must roll up our sleeves and start digging through these datasets.

Does that mean that all journalists need to sign up to Code Academy, take an evening course in programming or finally get around to memorising the parameters for Excel's VLOOKUP function? After all, while tools like Tableau or Statwing are really good at visualizing single sets of data, the data you need for your story is probably spread across multiple files, mixed in with all sorts of stuff you're not interested in and released separately for each year by each local authority. And the tools for people who want to stitch all that together and find the interesting numbers usually involve a bit of programming. Excel has its formulas and OpenRefine has its expression language. But for people who just don't think in that way, the tools for manipulating data are a little thin on the ground.

Until now that is.

Having spent most of my career as a programmer seeing people wrestling with spreadsheets and taking days to do things I knew the computer could do in seconds, I started to wonder if there was a more visual and interactive way to work with data. It was that pondering which eventually led to QueryTree, a drag and drop tool for working with data.

My company created QueryTree to be a simple and visual environment in which tools that manipulate data can all work together to give the user a lot of power and flexibility. QueryTree makes it easier to apply a series of changes to a set of data by representing each step in the process as an icon (referred to as a “tool”) on a graphical design surface. Each icon or tool represents a function such as sorting, filtering, selecting and grouping. Each tool takes in some data, or gives out some data, or both. Each tool has some settings and displays its results in the result area on the bottom half of the screen. And by sticking to those simple rules, each tool can work with all the other tools, but still stay very focused on the one job that it does. As a user, you can connect a number of tools together in a chain to achieve quite complicated results, or just learn the one or two tools that you need and ignore the rest. 

For example, say you were investigating violent crime but the only data available is a CSV containing details of all reported crimes for a particular year. You'll want to filter out the non-violent crimes, then group them up by type, counting the total number of incidents for each, and then sort from highest to lowest before exporting into another tool. Each step of that process would be represented in QueryTree as a separate tool, each one taking the output from the previous step and changing it in some way. So, our example would look like this:

DDJ_-_What_Is_QueryTree_-_Image1.png

To build this, a user drags and drops each tool into place from the toolbar. Each tool does one relatively easy to understand job like loading from a file, filtering or grouping. Clicking on each tool in the chain displays the data set at that point in the process in the results area.

However, investigations into data are not usually that linear. During the course of your digging you may want to take a look at all the violent crimes on a map to see if anything jumps out at you, or merge several other datasets together. You may want to try plotting various graphs to get a sense of what the data contains. In fact, your QueryTree worksheet would probably look more like this:

DDJ_-_What_Is_QueryTree_-_Image2.png

Other tools tend to only present one version of the data at any one time. If you sort your table in Excel, then your table is now sorted. This is a problem because going back and forth, comparing different views, or just explaining to your editor what you've done, requires you to keep a record of each step. With QueryTree not only is each tool simple and code free, but each tool also creates a separate version of the data, and the path the data has taken is laid out for all to see. 

QueryTree is free to try and can be accessed at querytreeapp.com. Over the coming months we'll be working on more tools for loading in different types of data, for cleaning data and for creating more impressive visualisations. If you'd like to hear about new features and updates you can follow us on Twitter at @QueryTreeApp, or sign up to our newsletter.

]]>
2013-06-14T02:29:04+00:00
The ProPublica Nerd Blog Presents: Design Principles for News Apps & Graphics http://datadrivenjournalism.net/resources/the_propublica_nerd_blog_presents_design_principles_for_news_apps_graphics http://datadrivenjournalism.net/resources/the_propublica_nerd_blog_presents_design_principles_for_news_apps_graphics#When:07:54:51Z

Originally published by Lena Groeger on ProPublica on 3 May 2013 under a Creative Commons license.

Screen_shot_2013-06-04_at_10.49.50_AM.png

While the tools and techniques to present large datasets in graphics and news apps may differ from project to project, the basic design principles stay pretty much the same. Many of these principles should seem pretty familiar – even if you’ve never studied design formally, you probably know some of them instinctively. Let’s name them, explain why they work and see how other designers, especially news designers, use them. Once you recognize the concepts, you’ll become more conscious of when and how to use them in your own projects.

Principle 1: Invisibility – Good Design Disappears

Invisibility isn’t a design principle per se, but for me it’s a helpful way to think about what these basic design principles are actually for.

Invisibility means design that’s clear, design that goes unnoticed (except perhaps by other designers) and design that “just works.” When a design presents distractions, discrepancies, conflicts and mismatches between what we expect and what we see, our brains have to put in more effort to figure out what’s going on. Good design eliminates as much of this extra effort as possible. Through placement and grouping and hierarchy, good design frees up mental space so users can think about content, and not where they’re supposed to be looking and how to interpret what they’re seeing.

Invisible design is strong design – if people don’t notice it, it’s not getting in the way.

Principle 2: Scale – Show the Near and the Far

Years ago a design instructor explained to me that whenever he designed a poster, he kept two viewpoints in mind. First, the viewpoint of the person seeing the poster from across the street, who could only make out the large forms and main ideas. Second, the viewpoint of the person who had crossed the street and was now looking at the poster close up, who could see all the details and wanted to find specific information. As a student, I thought about this a lot in the context of making posters, but didn’t realize how much value it had in the news world.

So I was surprised when, a few months after I began working at ProPublica, we talked about some of the design theories behind our work. Our news apps have (at least) two views. First, a “far” view, which is the national picture, how places compare, why you should care about this story as a whole; and a “near” view, which is, ideally, readers’ own personal stories – their town, their hospital, their school, etc. – all of the details that matter to them personally.

Screen_shot_2013-06-04_at_10.53.51_AM.png
The far view (left) and the near view (right). ProPublica, The Opportunity Gap

The “near and the far” serves not only as a metaphor for organizing a news app, but also as a guide for organizing a single page. Users will subconsciously evaluate each element’s importance based on its position and size relative to all the other elements on the page, so the 72pt text at the top will seem more important than the 9pt fine print in the lower-left-hand corner. These attributes establish a visual hierarchy, and let users know where to focus their attention first. It’s the designer’s job to make sure the hierarchy of elements matches the editorial goals. A good start is to make the most important editorial element (that is, “the nut” of the story) the most noticeable thing on the page.

The editorial and design considerations that go into poster design apply in much the same way to making a news app. Both require an understanding of the principle of scale. Whether it’s across the street vs. close-up or the national picture vs. the personal, scale works by grabbing people’s attention and interest, giving them context, and then drawing them in closer and showing them the details.

Principle 3: Alignment – Don’t Be Haphazard

In graphic design you often hear about alignment in terms designing on a grid, or making text flush left or right. These are examples of a higher principle: Everything should be positioned for a reason, and should align with related elements. Strong lines can help organize elements, so proper alignment gives a sense of structure and unity. When in doubt, find a line and stick with it.

Alignment is the driving force behind the table, one of the most common ways to present data. So prevalent we probably pass them over most of the time, tables are great examples of how alignment can aid the presentation of a lot of data. NPR’s Big Board was just such a design that managed, through alignment, to present a lot of information very clearly and intuitively about the 2012 election (their main election page also made great use of tables).

Screen_shot_2013-06-04_at_10.56.30_AM.png
Tables: pretty good at alignment. NPR, Elections Big Board

Alignment can also be used to represent more abstract concepts. This NY Times graphic on women in the Senate uses alignment to organize time. In this case, a line’s placement on the horizontal axis indicates where it lies in time.

Screen_shot_2013-06-04_at_10.57.53_AM.png
Alignment to indicate place in time. New York Times, Women in the Senate

When elements are arranged along a strong line, users don’t need to be able to see the line to know that it’s there. In many cases you can safely eliminate table borders. Strong lines speak for themselves – even when they’re invisible.

Careful alignment also makes it easy for the user to detect outliers, because anything that deviates from the norm will immediately stand out. The humble bar chart is a great example of this phenomenon. By aligning all bars along the bottom, our brain can rapidly detect length differences at the top.

A centered alignment has a weaker line, so a good rule of thumb is to avoid centering things. If you’re going to center it, do it consciously, not by default.

Principle 4: Repetition – Repeat Elements to Emphasize Common Form or Function

Repeating visual elements like shapes, typefaces and colors establishes a sense of purposefulness and internal consistency. You can use these elements to unify a design across multiple pages or products, so that users know they belong together. For example, the science magazine Nautilus uses a particular color yellow across its website (in bits of highlighted text or section backgrounds), to establish a cohesive look and feel and remind you that all pages are part of the same publication. You can also use repetition to associate elements in a user's mind with a particular function, such as the ability to input text or sort a table. For example, a blinking cursor is repeated all over the web to mean “if you type the text will appear here.”

Repetition also helps the user instantly see what is similar, and reveals what is different. By seeing the same form repeated, the user can easily recognize slight variations. This is the idea behind small multiples, which are sequences of small graphics that display differences in data and can be easily compared. Small multiples have the added benefit of not relying on the viewer’s memory to see the comparisons, because every element is presented to you at the same time. In this NY Times graphic, you can actually see changes in drought patterns because your eye becomes accustomed to the small U.S. maps and perceives only the changes between them.

Screen_shot_2013-06-04_at_11.00.56_AM.png
Small-multiple mini maps make comparisons easy. New York Times, Drought’s Footprint

In the same way, repetition can be useful to establish a visual cue for the function of an element across projects. At ProPublica we have an internal set of standard design elements, such as a particular color blue that we use for all the search boxes or input fields, which we call “do something blue.” We repeat this color throughout our site so that people who’ve used our apps more than once begin to associate the color with the ability to search, filter or participate.

Screen_shot_2013-06-04_at_11.03.55_AM.png
“Do Something Blue” in the wild. ProPublica, Dialysis Tracker, The Opportunity Gap, Dollars for Docs

In many cases you don’t have to create these visual cues from scratch. The web is filled with design conventions that you should follow as much as possible so that your design is intuitive.

Principle 5: Contrast – Don’t Be a Wimp*

Use design choices like color, size, and typography boldly to get people to notice what you want them to. The eye is drawn to movement on a still page, bright colors on a page of muted colors and bold elements on a page of neutral ones. Contrast attracts the eye, provokes emotion and directs attention. So don’t be shy about it – too much subtlety will make your users think the element is broken or its design a mistake. If a user has to work to figure out whether a switch from 12-point type to 13-point type was actually on purpose (“is it really different? why is it different? what does it mean?!”), then she is working too hard and has missed your point. Make it obvious.

For this graphic on drones, we wanted to point out contradictions between statements from officials on the CIA drone program. Among the long list of statements, we used a contrasting color to indicate the outliers, highlighting the contradictory statements in red. Strong color differences are processed very rapidly by the brain’s visual system. They immediately pop out from the background, which in this case was the effect we intended.

Screen_shot_2013-06-04_at_11.08.59_AM.png
Contrast to highlight contradictions. ProPublica, Stacking Up the Administration’s Drone Claims

Contrast can reveal the variability that exists in the data, add visual interest and make a point. It can also tell a story. This NY Times interactive on the aftermath of a deadly tornado lets you look simultaneously at a street in Moore, Okla. before and after the storm. Here the contrast comes from the photos themselves. Our brains are attuned to notice the differences, so highlighting them is a matter of capturing and closely coupling the panoramas.

Screen_shot_2013-06-04_at_11.13.10_AM.png
Contrast to compare then and now. New York Times, Before and After: 360° Views From Moore, Okla.

If interesting differences are the “story” behind your graphic, then make them clear and very obvious. Design consistency shouldn’t get in the way here. In our drones example, if the contradictory statements were just bolded or slightly bigger, they may not have had the same effect.

*Shamelessly stolen from The Non Designers Design Book

Principle 6: Proximity – Keep Things Together That Belong Together

Proximity is like organizing a drawer. Things that are similar, like socks, get put together. Shirts go someplace else. If related things are grouped together, they are easier to find, and no one is forced to hunt around for information (or socks).

Putting things together visually that belong together conceptually not only reduces clutter, but helps establish a visual hierarchy as you group and subgroup different elements into discrete visual units. Fight the inclination to evenly spread out all your content and fill every pixel of the screen. Well-organized empty space can direct people’s attention in the same way a headline can. Don’t fear white space.

The HuffPost’s House outlook page uses proximity to group related elements (the half circle representation of seats and five important numbers).

Screen_shot_2013-06-04_at_11.15.35_AM.png
Effective grouping and white space make for a clear presentation. Huffington Post, 2012 House Outlook

Proximity means keeping text close to the data it is describing, legends close to the map they’re explaining, etc. The “annotation layer,” a phrase coined by by Amanda Cox for guiding a user’s attention and providing clues to what’s important, also works because relevant text is positioned next to the elements being described. (Here’s a wonderful example of an annotation layer adding context and explanation.)

In this reconstruction of the Boston Marathon bombing, descriptions and labels are placed as close as possible to the relevant parts of the image (sometimes directly overlaid) instead of, say, listed all at the bottom. WNYC’s project on Fielder Avenue also keeps a navigational mini-map, photographs, and quotes together to emphasize their connection.

Screen_shot_2013-06-04_at_11.18.42_AM.png
Keep your annotations close. New York Times, Reconstructing the Scene of the Boston Marathon Bombing

Screen_shot_2013-06-04_at_11.19.35_AM.png
Photos, quotes and a map are grouped to create visual units. WNYC, Fielder Avenue

Principle 7: Intuitiveness – An Element’s Features Should Suggest Its Function

In the physical world, some objects are better suited for some things than for others. A door handle is better for pulling than a flat piece of metal. A wheel is better at rolling than a square box. The properties that make an object good at its function are called affordances (as in, a wheel affords rolling.) When an object is designed well, its affordances match its intended function. We know to pull the door open without any instruction, because its handle affords pulling.

In the virtual realm, we don’t have too much control over affordances. You can click every inch of your screen (the screen affords clicking) even if nothing happens in response. What we do have control over are “perceived affordances,” or what the user perceives the intended function to be.

It turns out that we bring plenty of intuitions and learned conventions about what affordances should look and work like, and about how to read and navigate visual information. Some of these are simply cultural conventions, for example that red means “stop” and green means “go,” or that scroll bars are located on the right side of the screen. But they’re very real. Put the scroll bar on the left and things will seem broken.

Other conventions are more deeply rooted in how our brains work. For example, take line and bar charts. We tend to read lines as showing trends – they are continuous, they connect, they show things that have a common dimension. On the other hand we read bars as showing discrete comparisons – they are categorical, they separate, they present things that are distinct. Those intuitions are so strong that they influence the way we interpret data, even if it leads to rather curious interpretations.

A couple years ago researchers had students look at either line or bar charts representing the same data: the height of men and women. One student who saw the line chart version explained, “as you get more male, you get taller.” No kidding. Our intuition that lines show trends is so strong that it biases interpretations in unexpected (and absurd) ways. Don’t fight these intuitions. Work with them.

In The Design of Everyday Things, Donald Normal writes: When a device as simple as a door has to come with an instruction manual – even a one-word manual – then it is a failure.

A well designed news application or graphic needs no instruction manual for the majority of its users. Understand your users' intuitive sense of things and don’t violate them by mistake. Pick chart types that make sense for the data you’ve got. Avoid using a line chart to display categorical data or bar charts to show trends or continuous data. In a similar vein, keep consistent with web standards like blue links and right-hand scroll bars and buttons that look like buttons and a mouse pointer that turns into a hand when hovering over something clickable. On tablets, follow the conventions of that platform. They’re there to help you.

When you break these intuitions and conventions, do it purposefully (be obvious!) and know you might have to give people clues on how to use your design.

Stop it.

Principle 8: Simplicity – Take Out What’s Extra And Be Prepared to Kill Your Darlings

If something isn’t necessary, take it out. Be ruthless in the cause of clarity and in defense of your user’s attention. Lose extraneous visual elements like gratuitous decorations, and make editorial decisions to eliminate anything that isn’t crucial to what you are trying to present. Chop off excess explanation and over-long copy, and eliminate user-pathways that don’t tell a story to your users. Not every column of the data you’ve analyzed needs to be in your app or graphic, and not every finding of your reporting needs to be in your story. Design is just as much about what you leave out as what you put in.

This doesn’t mean that you should ignore delight and surprise – the most minimalist, highest data-ink ratio is not always the right option. Projects are made with different goals for different audiences. But all elements should be there for reason.

Now Your Turn

I hope that these basic design principles will come in handy the next time you want to present lots of information visually. Try to apply each principle in turn, making adjustments to sizing or contrast, grouping elements together, taking out what’s unnecessary and making sure intended functions are intuitive. If something isn’t working, put the problem into words, and try to identify a principle that might fix it. Soon you’ll recognize these principles everywhere, your designs will be more deliberate, and you’ll be more in control. Good luck!

Inspiration, Resources and Things You Should Really Read Now

The Non-Designers Design Book, Robin Williams (designer, not actor). A fantastic intro to graphic design.

The Design of Everyday Things, Donald Norman. You will never look at doors or teapots the same way.

The Tufte Classics, Edward Tufte. If you’ve got these sitting on your bookshelf looking intimidating, open them up! They’re mostly annotated pictures and full of great examples.

The Functional Art, Alberto Cairo. A good introduction to information graphics, from a journalist’s perspective.

Universal Principles of Design, William Lidwell. An almanac of design principles, from graphic to interaction to architecture.

Visualizing Thought (paper), Barbara Tversky. A detailed look at how we interpret visual expressions like symbols and marks.

For something a bit more technical, we’ve put together a news apps style guide of best practices.

]]>
2013-06-06T07:54:51+00:00
Groupthink: Grouping Values in Excel Pivot Tables http://datadrivenjournalism.net/resources/groupthink_grouping_values_in_excel_pivot_tables http://datadrivenjournalism.net/resources/groupthink_grouping_values_in_excel_pivot_tables#When:23:09:42Z By Abbott Katz, London-based Excel instructor and freelance writer, author of Excel 2010 Made Simple and spreadsheetjournalism.com.

Mapping data is all the rage these days, a species of digital journalism founded atop columns of paired, spreadsheet-stamped latitudes and longitudes (lats and longs for short) that pump location-specific event information (e.g., records of crimes, bus stops) into Tableau-like apps that pin them to the maps. In fact, latitudes and longitudes can be thought of as planetary cell addresses.

In fact lats and longs can be harnessed into faux, pivot-table-plotted maps that do a creditable job of putting the data in their place even absent third-party intercession (see also my October 25 and November 1 spreadsheetjournalism posts). But quite apart from the data viz side of the question, there may be good reasons to place latitudes and longitudes in the service of more standard spreadsheet functionalities.

For example, one could break out crime distributions by the lats and longs in which they are perpetrated – but a problem besets that intention. Because latitudes and longitudes are often extended to many decimal points, their very precision imposes what could be termed a hyper-granularity upon the data – there's too much precision for the picture to cohere. Put a photograph under a microscope and all you see is pixels; it takes a prudent distancing before you get to see the picture for what it is. Thus if you're running the numbers through a pivot table you'll probably need to group the lats and longs into workable intervals. 

Here, for example, is an excerpt of some pivot-tabled latitudes in Washington DC, gathered for the purposes of pinpointing crimes by lats and longs:

Screen_shot_2013-05-28_at_7.01.30_PM.png

By clicking PivotTable Tools > Options > Group Selection in the Group button group, you'll open the Grouping window:

Screen_shot_2013-05-28_at_7.02.48_PM.png

Which allows you to identify a "By" interval, by which the latitudes can be grouped.  If I decide to group these by say, .01 of a degree (remember we're looking at data emanating from but one city), I get something like this:

Screen_shot_2013-05-28_at_7.03.33_PM.png

And that's rather dizzying, not the sort of thing you'll want to inflict upon a readership – not if you want them to keep reading. What's more, the numbers you see above can't be reformatted; because once they're grouped they acquire the status of inert text, and nothing more can be done with them.

What to do, then? Well, the standard workaround is to return to the source data, insert two columns (one each for latitude and longitude) and call upon the ROUND function. ROUND is a straightforward expression that empowers you to round a number to a desired level of precision – or redefined in the terms of our discussion, achieves a de facto grouping of values.

For example, if cell A3 contains this latitude:

37.76990984

Then =ROUND(A3,2) will return

37.77

in the cell in which it was composed, having rounded the value to two decimal points – and keep in mind that ROUND does not deliver the formulaic equivalent of repeated clicks of the Decrease Decimal button on the Home ribbon. That latter option only formats the number's appearance, and continues to honor its primeval quantitative value – that is, 37 plus those eight decimal points. ROUND, on the other hand, actually changes the value of its result – in this case to a "real" 37.77. (Note that ROUND also does its work on integers as well. =ROUND(23567,-2) - and note the negative number in the expression – smooths that value to 23600, and =ROUND(23567,-3) returns 24000.)

Thus if one implements the ROUND strategy and introduces the rounded data to the pivot table instead, the latitudes would read

Screen_shot_2013-05-28_at_7.05.29_PM.png

A far more palatable reading of the numbers. The above data denote value thresholds, i.e., latitudes falling between 37.79 and 38.75 will be grouped by the 37.79 value, etc.

Moreover, the numbers above are in fact numbers, and as such are amenable to genuine numeric formatting. Thus the latitudes could be formatted to two decimal places, and as such would turn that 38.9 into a decimal-consistent 38.90. Do the same for longitudes and the Washington DC crime data begin to look something like this:

Screen_shot_2013-05-28_at_7.06.20_PM.png

Now that's all very good, but sometimes one's data demands can't be satisfied by entrusting the data to ROUND alone. Consider this example:  you're looking at US presidential election data and want to learn something about vote distributions by county size, say in intervals of 25,000 voters. Now of course you could subject the data to Excel's Grouping mechanism, but again, that decision might or might not be visually fetching;

Screen_shot_2013-05-28_at_7.08.26_PM.png

That's still all a bit too dense for my taste, and remember, those numbers aren't numbers – they're text, and can't be reformatted. But ROUND won't work here either, because that function only coarsens a value upwards – e.g., 525064 could be rounded to 525100 or even 530000, but you can't size it to the nearest interval of 25000.

But there is an alternative, really a pair of them, neither of which I was unaware until a short while ago – FLOOR and CEILING. These sibling functions round a value to the nearest specified multiple, either below or above the value.

What does that mean? It's really quite simple. If I enter

=FLOOR(525064,25000)

I realize 525000 – the closest value to 525064 that's both divisible by 25000 and lower than the original 525064. Enter

=CEILING(525064,25000)

And you return 550000 – the closest value to 525064 divisible by 25000 that's higher than the original number.

Thus if I want to break out the above county voter totals (and assuming they're in say, column B) by multiples of 25000 I can insert a new column among the data and enter

=CEILING(B2,25000)

And copy it down the column. Thrust those data into a pivot table and I get

Screen_shot_2013-05-28_at_7.09.49_PM.png

Where each Row Label value offers itself as the upper limit of any voter total threshold. And again, because we've retained the values' numerical status we can reformat some commas into the mix (and perhaps right-align them all, too):

Screen_shot_2013-05-28_at_7.10.32_PM.png

Now isn't that easier to read?

]]>
2013-06-03T23:09:42+00:00