Safari Books Online is a digital library providing on-demand subscription access to thousands of learning resources.
This chapter is a modest attempt to introduce Natural Language Processing (NLP) and apply it to the unstructured data in blogs. In the spirit of the prior chapters, it attempts to present the minimal level of detail required to empower you with a solid general understanding of an inherently complex topic, while also providing enough of a technical drill-down that you’ll be able to immediately get to work mining some data. Although we’ve been regularly cutting corners and taking a Pareto-like approach—giving you the crucial 20% of the skills that you can use to do 80% of the work—the corners we’ll cut in this chapter are especially pronounced because NLP is just that complex. No chapter out of any book—or any small multi....