Creating a word cloud on R-bloggers posts
This post will go through how to create a word cloud of article titles scraped from the awesome R-bloggers. Our goal will be to use R's rvest package to search through 50 successive pages on the site for article titles. The stringr and tm packages will be used for string cleaning and for creating a term document frequency matrix (with tm). We will then create a word cloud based off the words comprising these titles. First, we'll load the packages we need. Let's write a function that will take a webpage as input and return all the scraped article titles. The above function takes an input, called site, which will be the URL of a specific webpage on R-bloggers. We then use rvest's read_html function to scrape the HTML from…