Stories filed under: "faked"

Faux Randomness Strikes Again: How Researchers Realized Research 2000's Daily Kos Data Looked Faked

from the random-ain't-so-random dept

Tue, Jun 29th 2010 02:52pm - Mike Masnick

You may have heard by now that the political website Daily Kos has come out and explained that it believes the polling firm it has used for a while, Research 2000, was faking its data. While it’s nice to see a publication come right out and bluntly admit that it had relied on data that it now believes was not legit, what’s fascinating if you’re a stats geek is how a team of stats geeks figured out there were problems with the data. As any good stats nerd knows, the concept of “randomness” isn’t quite as random as some people think, which is why faking randomness almost always leads to tell-tale signs that the data was faked or manipulated. For example, one very, very common test is to use Benford’s Law to look at the first digit of data in a data set, because in a truly random set, the distribution is not what people usually expect.

In this case, the three guys who had problems with the data (Mark Grebner, Michael Weissman, and Jonathan Weissman) zeroed in on just a few clues that the data was faked or manipulated. The first thing they noticed was that when R2K did polls that tested how men and women viewed certain politicians or political parties (favorable/unfavorable) there was an odd pattern: if the percentage of men that rated a particular politician favorable or unfavorable was an even number, so was the the percentage of female raters. It seemed like these two points always matched up. If the male percentage was even the female percentage was even. If the male percentage was odd, the female percentage was odd. Yet, as you should know, these are independent variables, not influenced by each other. That 34% of men find a particular politician favorable should have no bearing on why an even percentage of women find that politician favorable. In fact, this happened in almost every such poll that R2K did, to such a level as to suggest it being as close to impossible as you can imagine:

Common sense says that that result is highly unlikely, but it helps to do a more precise calculation. Since the odds of getting a match each time are essentially 50%, the odds of getting 776/778 matches are just like those of getting 776 heads on 778 tosses of a fair coin. Results that extreme happen less than one time in 10²²⁸. That’s one followed by 228 zeros. (The number of atoms within our cosmic horizon is something like 1 followed by 80 zeros.) For the Unf, the odds are less than one in 10²³¹. (Having some Undecideds makes Fav and Unf nearly independent, so these are two separate wildly unlikely events.)

There is no remotely realistic way that a simple tabulation and subsequent rounding of the results for M’s and F’s could possibly show that detailed similarity. Therefore the numbers on these two separate groups were not generated just by independently polling them.

The other statistical analysis that I found fascinating was that when you looked at weekly changes in favorability ratings, the R2K data almost always changed a bit. But, if you look at other data, no change is the most common result. As they point out, if you look at, say, Gallup data, you get this nice typical bell curve:

But if you look at the R2K data, you get things like the following:

Notice that substantial dip at the 0% mark. That seems to indicate a likelihood of faked or manipulated data, from someone who thinks that “random” data means the data has to keep changing. Back to Grebner, Weissman and Weissman:

How do we know that the real data couldn’t possibly have many changes of +1% or -1% but few changes of 0%? Let’s make an imaginative leap and say that, for some inexplicable reason, the actual changes in the population’s opinion were always exactly +1% or -1%, equally likely. Since real polls would have substantial sampling error (about +/-2% in the week-to-week numbers even in the first 60 weeks, more later) the distribution of weekly changes in the poll results would be smeared out, with slightly more ending up rounding to 0% than to -1% or +1%. No real results could show a sharp hole at 0%, barring yet another wildly unlikely accident.

Kos is apparently planning legal action, and so far R2K hasn’t responded in much detail other than to claim that its polls were conducted properly. I’m not all that interested in that part of the discussion however. I just find it neat how the “faux randomness” may have exposed the problems with the data.

Filed Under: data, faked, random
Companies: dailykos, research 2000

35 Comments

Expand

Follow Techdirt

Subscribe to Our Newsletter

Essential Reading

The Techdirt Greenhouse

Read the latest posts:

Read All »

Trending Posts

Techdirt Insider Discord

The latest chatter on the Techdirt Insider Discord channel...

Older Stuff

Tuesday
13:30	Techdirt Podcast Episode 449: The Dangers Of Product Design Liability For Social Media (0)
10:52	The New York Times Got Played By A Telehealth Scam And Called It The Future Of AI (35)
10:47	Daily Deal: Costco 1-Year Gold Star Membership + $20 Digital Costco Shop Card (2)
09:24	Trump's Office Of Legal Counsel Says Trump Doesn't Need To Follow The Presidential Records Rules (10)
05:25	America First? Paramount Finalizes $24 Billion In Middle East Backing For Warner Bros Deal (8)
Monday
20:10	Congress Wants To Put The Law Behind A Paywall. Again. (16)
15:10	Trump Fires Attorney General Pam Bondi For Not Making His Vindictive Fantasies A Reality (20)
13:06	UK Politicians Continue To Miss The Point In Latest Social Media Ban Proposal (6)
11:02	Trump Celebrates Easter By Dropping An F-Bomb, Threatening More War Crimes (61)
10:56	Daily Deal: The Academy of Game Art Bundle (0)
09:25	Jacob Siegel's Error-Filled Book On 'Censorship' Got Fact-Checked. He's Calling It Censorship. (24)
05:24	Supreme Court Shrugs Off Opportunity To Save The First Amendment From The Fifth Circuit's Antipathy (6)
Sunday
12:00	Funniest/Most Insightful Comments Of The Week At Techdirt (12)
Saturday
12:00	Game Jam Winner Spotlight: CARAMENTRAN (3)
Friday
19:39	Minnesota Kicks Off Legal Battle With Trump Administration To Hold ICE Shooters Accountable (14)
15:32	In Chiles V. Salazar The Supreme Court Issues A Bad Good First Amendment Decision (26)
13:08	Can Agentic AI Coding Tools Finally End Copyright For Software While Re-Inventing Open Source? (49)
13:03	Daily Deal: Hypergear 3-in-1 Wireless Charging Dock (0)
11:02	The Social Media Addiction Verdicts Are Built On A Scientific Premise That Experts Keep Telling Us Is Wrong (24)
09:24	Senators Ask Tulsi Gabbard To Tell Americans That VPN Use Might Subject Them To Domestic Surveillance (11)
05:24	The Trump Administration Is Trying To Steal $21 BIllion Earmarked For Better Broadband (8)
Thursday
20:08	DOGE Goes Nuclear: How Trump Invited Silicon Valley Into America’s Nuclear Power Regulator (9)
15:20	Ctrl-Alt-Speech: Age Old Questions (0)
13:01	The AI Doc’s Falsehoods And False Balance (27)
11:04	Meta Caves To The MPAA Over Instagram's Use Of 'PG-13,' Ending A Dispute That Was Silly From The Start (13)
10:59	Daily Deal: Opusonix Pro Subscription (0)
09:37	Trump's Anti-Migrant Surge Is Now A Mudslide That's Wiping Out What's Left Of His DOJ (17)
05:32	WSJ: Lobbyists Easily Destroyed Any Semi-Serious Antitrust Enforcers Left In MAGA (4)
Wednesday
20:14	Federal Cyber Experts Thought Microsoft’s Cloud Was “A Pile Of Shit.” They Approved It Anyway. (12)
15:36	South Dakota GOP, Governor Get Their Voter Suppression On (7)

Faux Randomness Strikes Again: How Researchers Realized Research 2000's Daily Kos Data Looked Faked

from the random-ain't-so-random dept

Get all our posts in your inbox with the Techdirt Daily Newsletter!

The Techdirt Greenhouse

Trending Posts

Tuesday

Monday

Sunday

Saturday

Friday

Thursday

Wednesday

More

Tools & Services

Company

Contact

More

from the random-ain't-so-random dept

Techdirt Daily Newsletter

Get all our posts in your inbox with the Techdirt Daily Newsletter!

The Techdirt Greenhouse

Trending Posts

Email This Story

Tools & Services

Company

Contact

More