Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Tuesday, September 29, 2009

The Flu Shot Controversy

A large Canadian study has found that people who got the seasonal flu shot last year were more likely to get H1N1 this spring than people who didn't get the shot. The report is not yet peer-reviewed, but reports of the report are causing Canadian officials to rethink their flu shot policy. However, the researchers should have expected that certain biases would provide this result, without implying that the seasonal flu shot somehow causes H1N1:

1. People who get the seasonal flu shot are more likely to get flu
Some people never get the flu, because of natural immunity or because they don't spend time in close proximity to infected people. These people are much less likely to get flu shots.

Conversely, some people are much more prone to getting sick, due to poor health, close proximity to infected people, or whatever. These people are much more likely to get the flu shot.

So in the sample group, you would expect more people who got the flu shot to get this new strain of flu. That doesn't mean that the flu shot caused the flu.

2. People who get the flu shot are more likely to be diagnosed with H1N1
This past spring, most people who got sick with H1N1 got mild cases. Most were probably not even diagnosed. The people who were diagnosed with H1N1 are probably those people who go to their doctor more often. People who go to their doctors more often are more likely to get the flu shot. Therefore, of the people who got H1N1, those who are diagnosed are more likely to have had the flu shot.
--
Correlation does not imply causation.

This sort of research can suggest lines of laboratory research, but it means little without the research. And yes, that applies to all the other statistical health studies we read about in the newspaper.

Unfortunately, immunizations have become the hot topic of people who are paranoid about the medical profession: people who hate doctors, or mistrust science in general, or think BigPharma is manipulating health issues to boost profits.

As for myself, for most of my life I got the flu every year, usually getting sick as a dog and missing a full week of work or more. Every year since 1997 I've had a flu shot and in that time I haven't had the flu. Ontario has now decided to delay seasonal flu shots till after the H1N1 shots, and that increases my likelihood of getting the flu. So I feel I have a personal stake in this. For those who don't want to get flu shots - fine; as long as they stay away from other people when they get sick, they can do what they like. But as a public policy, flu shots work, and shouldn't be delayed until the middle of flu season on such a questionable and preliminary study.

###

Wednesday, July 22, 2009

Making Information Pop



As a technical writer, I am constantly thinking about how to help readers absorb and retain information. It's not enough for my documentation to be correct and complete; it has to be useful to readers in that it results in a more effective user experience. So the metric of quality for software documentation is how well users can use the software (in full realization that nobody likes to read the documentation and will consult it as little as possible). That means that I have to be very creative when creating help.

Re the above video, it's exciting to see someone who got creative about statistical reporting with such success. They have not only created more sophisticated analysis and a zippier view (using a Google gadget called motion charts), but he has turned a statistical lecture into storytelling.

I have long believed in the importance of an emotional element to help people absorb and retain information. I once read an abstract of a PhD thesis in technical writing that found people retain information better if it makes them happy - but I suspect happiness is just a subset of emotions that will achieve the same end. I recently saw a lecture (I've lost the link, unfortunately, but it was on TED) about how graphics engage different parts of the brain and get people to make connections that words alone cannot make.

Seeing this brilliant statistical presentation approach reinforces something I've thought for a long time: documentation should have graphics: cartoons and photos as well as UML diagrams and screenshots. As an example, I have an old copy of The Complete Idiot's Guide to Java 2, and it includes a repeated cartoon of a woman yelling out, "Hey stupid!" with a caption underneath that gives a little bit of info about Java. More than once I have flipped through the book reading all the "Hey stupid"s with rapt attention.

On one hand the field of technical writing is prone to a lot of gimmicky ideas and useless bells-and-whistles. On the other, it is held back by localization concerns: many cultures are said to object to representations of humans (even of hands), and Germans are said not to appreciate humor in business, and generally we are warned not to try anything funny in case it offends someone. Nevertheless, I think graphics are integral to engaging readers and should be considered more seriously as core elements of good documentation.

I specialize in documentation for developers, and mostly write reference material, tutorials, and very dense user guides. My books are not read cover to cover, but are generally read piecemeal as the developer needs the info. So not just any graphics will do: like Hans Rosling's statistical presentation software, above, we need to develop really effective, engaging graphics. Graphics that not just provide useful information, but that engage the mind in a novel way.

For more on the new approach to presenting statistics, and lots more examples of it, see gapminder.org.

Update: The lecture I was thinking of is Tom Wujec on 3 ways the brain creates meaning. The comments provide some needed caveats.

###

Friday, February 13, 2009

Risk Part 3: Case Study - How Poor Risk Management Caused the Crisis

The collapse of Lehman Brothers on September 15, 2008 is the event that tipped the financial system into full-blown crisis. Allowing Lehman to fail is arguably one of the worst decisions of the Bush administration. Why did Lehman fail? Here's an analysis of the problems at Lehman. This is an excerpt from an article by David Einhorn in the Global Association for Risk Professionals' publication GARP Risk Review. Keep in mind that this article was written after the spring of 2008 (when the financial crisis started) but before the collapase of Lehman, and that when it was written few people suspected any problems at Lehman:
The first question to ask is, how did this [financial crisis] happen? The answer is that the investment banks outmaneuvered the watchdogs, as I will explain in detail in a moment. As a result, with no one watching, the management teams at the investment banks did exactly what they were incentivized to do: maximize employee compensation. Investment banks pay out 50% of revenues as compensation. So, more leverage means more revenues, which means more compensation.
...
The second question is, how do the investment banks justify such thin capitalization ratios? And the answer is, in part, by relying on flawed risk models, most notably value at risk (VaR). VaR is an interesting concept. The idea is to tell how much a portfolio stands to make or lose 95% of the days or 99% of the days or what have you. Of course, if you are a risk manager, you should not be particularly concerned how much is at risk 95% or 99% of the time. You don’t need to have a lot of advanced math to know that the answer will always be a manageable amount that will not jeopardize the bank.

A risk manager’s job is to worry about whether the bank is putting itself at risk in the unusual times — or, in statistical terms, in the tails of distribution. Yet, VaR ignores what happens in the tails. It specifically cuts them off. A 99% VaR calculation does not evaluate what happens in the last 1%. This, in my view, makes VaR relatively useless as a risk management tool and potentially catastrophic when its use creates a false sense of security among senior managers and watchdogs. This is like an airbag that works all the time, except when you have a car accident.

By ignoring the tails, VaR creates an incentive to take excessive but remote risks. Consider an investment in a coin-flip. If you bet $100 on tails at even money, your VaR to a 99% threshold is $100, as you will lose that amount 50% of the time, which obviously is within the threshold. In this case, the VaR will equal the maximum loss.

Compare that to a bet where you offer 127 to 1 odds on $100 that heads won’t come up seven times in a row. You will win more than 99.2% of the time, which exceeds the 99% threshold. As a result, your 99% VaR is zero, even though you are exposed to a possible $12,700 loss. In other words, an investment bank wouldn’t have to put up any capital to make this bet. The math whizzes will say it is more complicated than that, but this is the basic idea.

Now we understand why investment banks held enormous portfolios of “super-senior triple A-rated” whatever. These securities had very small returns. However, the risk models said they had trivial VaR, because the possibility of credit loss was calculated to be beyond the VaR threshold. This meant that holding them required only a trivial amount of capital, and a small return over a trivial amount of capital can generate an almost infinite revenue-to-equity ratio. VaR-driven risk management encouraged accepting a lot of bets that amounted to accepting the risk that heads wouldn’t come up seven times in a row.

In the current crisis, it has turned out that the unlucky outcome was far more likely than the backtested models predicted. What is worse, the various supposedly remote risks that required trivial capital are highly correlated; you don’t just lose on one bad bet in this environment, you lose on many of them for the same reason. This is why in recent periods the investment banks had quarterly write-downs that were many times the firmwide modelled VaR.
...
Lehman’s management is charismatic and has almost cult-like status. It gets tremendously favorable press for everything from handling the 1998 crisis to supposedly hedging in this crisis to not playing bridge while the franchise implodes.

From a balance sheet and business mix perspective, Lehman is not that materially different from Bear Stearns. Lehman entered the crisis with a huge reliance on US fixed income, particularly mortgage origination and securitization. It is different from Bear in that it has greater exposure to commercial real estate and its asset management franchise did not blow up. Incidentally, neither Bear nor Lehman had enormous on-balance-sheet exposure to CDOs. At the end of November 2007, Lehman had Level 3 assets and total assets of about 2.4 times and 40 times its tangible common equity, respectively. Even so, at the end of January 2008, Lehman increased its dividend and authorized the repurchase of 19% of its shares. In the quarter ended in February, Lehman spent over $750 million on share repurchases, while growing assets by another $90 billion. I estimate Lehman’s ratio of assets to tangible common equity to have reached 44 times.

There is good reason to question Lehman’s fair value calculations. It has been particularly aggressive in transferring mortgage assets into Level 3. Last year, Lehman reported its Level 3 assets actually had $400 million of realized and unrealized gains. Lehman has more than 20% of its tangible common equity tied up in the debt and equity of a single private equity transaction — Archstone-Smith, a real estate investment trust (REIT) purchased at a high price at the end of the cycle. Lehman does not provide disclosure about its valuation, though most of the comparable company trading prices have fallen 20-30% since the deal was announced. The high leverage in the privatized Archstone-Smith would suggest the need for a multibillion-dollar write-down.

Lehman has additional large exposures to Alt-A mortgages, CMBS and below-investment-grade corporate debt. Our analysis of market transactions and how debt indices performed in the February quarter would suggest Lehman could have taken many billions more in write-downs than it did. Lehman has large exposure to commercial real estate. Lehman has potential legal liability for selling auction-rate securities to risk-averse investors as near cash equivalents.

What’s more, Lehman does not provide enough transparency for us even to hazard a guess as to how they have accounted for these items. It responds to requests for improved transparency grudgingly, and I suspect that greater transparency on these valuations would not inspire market confidence. Instead of addressing questions about its accounting and valuations, Lehman wants to shift the debate to where it is on stronger ground. It wants the market to focus on its liquidity. However, in my opinion, the proper debate should be about Lehman’s asset values, future earning capabilities and capital sufficiency.

In early April, Lehman raised $4 billion of new capital from investors, thereby spreading the eventual problems over a larger capital pool. Given the crisis, the regulators seem willing to turn a blind eye toward efforts to raise capital before recognizing large losses; this holds for a number of other troubled financial institutions. The problem with 44 times leverage is that if your assets fall by only a percent, you lose almost half the equity. Suddenly, 44 times leverage becomes 80 times leverage and confidence is lost. It is more practical to raise the new equity before showing the loss. Hopefully, the new investors understand what they are buying into, even though there probably isn’t much discussion of this dynamic in the offering memos. Some of the sovereign wealth funds that made these types of investment last year have come to regret them.

Lehman wants to concentrate on long investors; in fact, it went to great lengths to tell the market that it sold all of its recent convert issue to long-only investors. Putting aside the fact that some of the clearing firms have told us that this wasn’t entirely true, companies that fight short sellers in this manner have poor records. The same goes for companies that publicly ask the SEC to investigate short selling, as Lehman has done. There is good academic research to support my view on this point. As I have studied Lehman for each of the last three quarters, I have seen the company take smaller write-downs than one might expect. Each time, Lehman reported a modest profit and slightly exceeded analyst estimates that each time had been reduced just before the public announcement of the results. That Lehman has not reported a loss smells of performance smoothing. Given that Lehman hasn’t reported a loss to date, there is little reason to expect that it will any time soon. Even so, I believe that the outlook for Lehman’s stock is dim. Any deferred losses will likely create an earnings headwind going forward. As a result, in any forthcoming recovery, Lehman might underearn compared to peers that have been more aggressive in recognizing losses.

Further, I do expect the authorities to require the brokerdealers to de-lever. In my judgment, a back-of-the-envelope calculation of prudent reform would require 50-100% capital for no ready market investments; 8-12% capital for what the investment banks call “net assets”; 2% capital for the other assets on the balance sheet; and an additional charge that I don’t know how to quantify for derivative exposures and contingent commitments. Only tangible equity, not subordinated debt, should count as capital. On that basis, assuming that Level 3 assets are a good proxy for no ready market investments — assigning no charge for the derivative exposure or contingent commitments and assuming its asset valuations are fairly stated — Lehman, based on its November balance sheet, would need $55-$89 billion of tangible equity, which would be a three- to-five-fold increase.

See also:
Risk Part 1: Issues
Risk Part 2: The Mess
Risk Part 3: Case Study - How Poor Risk Management Caused the Crisis
Risk Part 4: Regulatory Revision
Risk Part 5: Capitalism 2.0
Risk Part 6: Moral Hazard
Risk Part 7: Some Basic Accounting Problems

Saturday, February 07, 2009

Measuring Risk Part 1: Issues

Risk is all the rage right now. I made a timely decision last year when I decided to move into financial risk statistics as a new area of expertise. The financial meltdown caught everyone unawares, and the idea of being better prepared next time has caught on big. Businesses are getting interested in finding more sophisticated ways to measure volatility and exposure, and new government regulations are also requiring they do so.

In the banks, risk is effectively a corporate governance function, not a business function. One reason for this is that risk is seen as an impediment to short-term profits (and hence bonuses). But it's also the case that people don't trust risk statistics. A lot of money goes into producing risk statistics, but people don't actually make decisions based on them... and when they do, their accuracy is so poor that they're hardly better predictors of financial gain than flinging a dart at a newspaper stuck to the wall.

Most risk measurement is based on normal distributions that simply don't exist in the stock market. A normal distribution can be assumed when you have a lot of small movement without a lot of outliers (meaning that severe market events would occur only once every few hundred years). In reality, we have severe market events every five years or so. Real distributions are not normal; they're skewed in various ways.

In addition, most measures of risk are measures of volatility around the mean. This assumes that upside volatility is as bad as downside volatility - hardly true! Plus, they assume that volatility and the correlations between assets are not affected by extreme market conditions, even though it has been shown that after extreme market shocks the volatility and correlations go haywire for a while.

But even if you use stable (non-normal) distributions, measure downside risk, and account for volatility clustering, there are basic limitations to fundamental analysis: how far can you go basing risk on historical market data? For example, you're not measuring the exposure of an asset to exchange rates: you're measuring the way the asset responded to exchange rate fluctuations in a particular historical period. You don't know why it fluctuated, so you can't predict it will follow the same pattern in the future. It's the fundamental problem of econometrics: correlation does not imply causation.

Even if the statistics were at all accurate, there are problems with how to use them. We need a more sophisticated vocabulary and set of statistics based on the purpose of the measurement, and we need a better understanding of how to apply the statistics to the real world. Regulators, corporations, risk managers and individual investors all have different needs for assessing risk and should in many cases use different statistics.

More to come in subsequent posts.

Risk Part 1: Issues
Risk Part 2: The Mess
Risk Part 3: Case Study - How Poor Risk Management Caused the Crisis
Risk Part 4: Regulatory Revision
Risk Part 5: Capitalism 2.0
Risk Part 6: Moral Hazard
Risk Part 7: Some Basic Accounting Problems

###

Sunday, February 11, 2007

Freakonomics

I read Freakonomics today. I was prepared to like it - it's just the sort of thing I like - but in the end I was disappointed and even angry at the authors.

The authors, Steven Levitt and Stephen Dubner, describe Freakonomics as "whatever freakish curiosities may occur to us." The book has little to do with economic analysis. It has little to do with analysis at all: analysis is used just so far as to provide a "Gee Wiz" type of answer, which is rarely considered deeply enough to provide real insight.

For example, the case study of three little girls (presumably around 8). The parents of one girl will let her go to visit one of her friends, who has a swimming pool, but not the other girl, whose parents have a gun. The authors provide the statistic that more children under 10 die of drowning in backyard pools than die in gun accidents so presto besto - the parents are irrational. But the parents are not dealing with average children under 10 - they're dealing with a specific set of circumstances. If most children who die in backyard pools are toddlers or older children who don't know how to swim, and if their 8-year old knows how to swim, then the swimming pool death statistic isn't relevant. For authors who claim to be trying to "understand the hidden side of everything", they seem to be more interested in taking cheap shots at the conventional wisdom than providing real insight.

Similarly with a chapter on how rich or educated people name their children vs poor or uneducated children: a 25-page chapter yields one insight, that the most popular names for poor or uneducated people are sometimes the names that rich or educated people named their children ten years before. That's it.

The chapter on crime is confused. Levitt's one original insight, that legalized abortion was a contributing factor to lower crime rates, is pretty neat. But the rest of the analysis is a mishmash of other theories, poorly and inconsistently explained. For example, they go back and forth on the importance of having two parents, and they don't even seem to notice they're doing it. There is no thought given to the different types of single-parent households: were the parents ever married; is the father around or paying support; has the child met the father. They make some startling assumptions, such as that sending more people to prison decreases crime rates, while Canadian studies show just the opposite - and at one point, they even argue the opposite (that increased jail time enabled gang members to get to know drug importers).

I like the repeated appeals to stop confusing correlation with causation. The book starts out well, and the first 100 pages are the best, but the whole thing is less than 200 pages, with the last half seeming a lot like filler. (My edition came with some "bonus material" - an overly laudatory bio of the authors and some previously-published articles that partly duplicate the material in the book.)

One of Levitt's recurring themes is that so-called experts often have their own agendas and so we should be wary of trusting them. That's an unintentionally self-referential argument. This book reads like a thrown-together rehash of old material, inadequately considered, with a hyped-up meaningless title, designed for one purpose - to make some quick cash. A more thoughtful and well-argued book might not have had the same mass appeal.

###

Friday, September 15, 2006

If You Had to Choose the Next Happy Meal, Which Meal Would You Choose?

This recent Gandalf poll seems a bit dodgy. The survey talked to 1,000 random Canadians. Here is the key question: "If you had to vote for the next Liberal leader, who among the candidates would you vote for?" The response was: Dryden 19%, Rae 17%, Ignatieff 10%, Dion 8%, Bennett 6%, Kennedy 4%.

So here's an equivalent kind of question: If you had to choose the next Happy Meal at Macdonalds, what would you choose? The options are: roasted vegetables on foccaccia; bacon, lettuce and tomato sandwich; deep-fried worms; or flaming bananas.

You might be a vegetarian who wouldn't be caught dead in Macdonalds but you support the concept of healthy alternatives so you choose the roasted vegetable sandwich, even though you don't plan to ever buy it. You might be an employee of Burger King and so you choose fried worms in the hopes of messing with the competition. You might not take any of this seriously and so think, Hey man, flaming bananas. Imagine all the kitchen fires! I'm going for that one.

So maybe our Macdonalds pollsters broke out the numbers for declared Macdonalds customers. Even then, the respondents might have the following thought processes: Geeze, foccaccia, never heard of it. I guess I have to choose the BLT, just because I know what it is (even though after an advertising campaign I'd realize that foccaccia is just a fancy name for white bread, and I'd prefer it). Or: This place is so stodgy, we need to shake it up. I don't care how; we just need change. And so I'm choosing flaming bananas, even though I'll never order them. Or: Wow, BLT, that would make a nice lunch; I'll pick that.

To the Gandalf question, "If you had to vote for the next Liberal leader, who among the candidates would you vote for?", people identified as Liberal voters answered: Dryden and Rae each 19%, Ignatieff 12%, Dion 8%, Kennedy 7%, Bennett 5%. Keep in mind though that only 30% of the respondents identified themselves as Liberals, so this data is based on 300 people across the country. Not enough.

Perhaps the biggest problem with "If you had to vote for the next Liberal leader, who among the candidates would you vote for?" is that it is ambiguous. Did respondents interpret it as, "If you had to choose the next Liberal leader, who would you choose?" or "If you had to vote Liberal in the next election, which leadership candidate would you rather vote for?" (Would you think it meant, "What do you think Macdonalds should put on its menu?" or "If you had to buy one of these meals, which one would you buy?")

Asking "if you had to choose..." tries to force people to come up with a response. A more meaningful question might be, "If the following items were on the menu, would you be likely to order them?" Or in the case of the Gandalf poll: "If Ken Dryden were leader of the LPC, would you vote Liberal?"

Way down at the end of the report the survey addresses this question. The question is "How likely would you be to vote for the LPC if it were led by the following candidates in the next election?" This is a fairly typical polling question that is designed to determine the pool of potential voters for each candidate (made up of respondents who say they are very likely, somewhat likely, or don't know whether they would vote Liberal if that person were leader).

In the Gandalf survey, the candidates for whom respondents said they were certain or likely to vote Liberal are: Rae 21%, Dryden 20%, Dion and Ignatieff each 16%, Bennett 13%, Kennedy 12%, Brison 11%, Findlay 10%. But the numbers also reveal the following:

* In another section of the survey, 70% of respondents said that they'd vote for a party other than the Liberals, but in response to this question, no more than 42% said they'd be unlikely to vote Liberal if any of the candidates were leader.
* In another section of the survey, 30% of the respondents said they'd vote Liberal, but in response to this question, at most 8% said they were certain to vote Liberal if a specific candidate were leader.

This shows incredible softness in opinion on both Liberal support and non-Liberal support. It might be useful data in a trend (Decima is guaging the potential voter pool on a repeating basis to show just that) but it's not clear how useful it is as a snapshot.

The report makes some claims that may be beyond its sampling. For example, it says, "In Quebec, Conservatives are now in a fight to hold their seats, and could lose up to seven of them to the BQ. BQ could come out of an election with 60 or more seats." This survey sampled 1,000 Canadians. If they made sure that the number of responses from each province was proportionate to the electorate, then about 250 of those surveyed were in Quebec. This seems like a small sample to forecast the outcome of over 60 ridings. To say anything about the outcome of any riding, I'd like to see a much better sample.

I'm very interested to see the results of the leadership convention delegate selections, which are due out in October. Until then I'd put surveys like this in the "yappa ding ding" category. Yappa ding ding is a Garifuna term meaning "something worth less than nothing".

###