Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> My approach when developing a predictive model has been always to throw the kitchen sink into a stepwise regression and then eliminate parameters based on their F Values. Is there a better way to do variable selection?

It depends on whether you care more about good predictions on data drawn from the same source or more about unbiased parameter estimates. For example, if you unwittingly add variables into your model that represent an intermediate outcome, you'll get selection bias and your parameter estimates will be off.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: