Chapter 2 Exploring the Network
2.1 Data collection strategy
At the time, it was possible to
The dataset is structured with six columns, each providing specific details about U.S. politicians:
nameThis column represents the full legal name of the politician. It is a nominal attribute.screen_name: This column indicates the Twitter handle associated with each politician. This attribute is particularly useful for social network analyses that involve online interactions.class: A categorical variable used to classify politicians based on their role in the government. In the sample provided, ‘S’ most likely signifies the Senate.role, This column specifies the particular governmental role held by the politician, such as a Senate position. This is a categorical attribute.party: This column designates the political party to which the politician belongs. In the U.S. context, ‘D’ indicates Democrats and ‘R’ indicates Republicans. This is a binary categorical attribute in the sample.state*: This column states the U.S. state that the politician represents. This is a nominal attribute and can vary depending on the scope of the politician’s role.- ```territoryw in the dataset represents an individual politician and offers a comprehensive set of variables that can be leveraged for various types of research analyses. These can include, but are not limited to, network analysis, sentiment analysis, and demographic studies. The dataset is thus well-suited for both univariate and multivariate analyses.
It also includes a table usp2020_edges with the following columns:
source, the screen_name of the sourcetarget, the screen_name of the target
First of all, we need to load the usp2020 dataset. This dataset contains the information about the individual politicians. We will use this dataset to create the nodes of our network.
## name screen_name class role party state
## 1 doug jones sendougjones S senate D alabama
## 2 lisa murkowski lisamurkowski S senate R alaska
## 3 kyrsten sinema senatorsinema S senate D arizona
## 4 martha mcsally senmcsallyaz S senate R arizona
## 5 john boozman johnboozman S senate R arkansas
## 6 tom cotton sentomcotton S senate R arkansas
Second, we need to load the usp2020_edges dataset. This dataset contains the information about the following relationships between the politicians. We will use this dataset to create the edges of our network.
## V1 V2
## 1 realdonaldtrump jim_jordan
## 2 realdonaldtrump vp
## 3 realdonaldtrump mike_pence
## 4 potus ustraderep
## 5 potus realdonaldtrump
## 6 potus stevenmnuchin1
Now, we can create the network. In order to do that, we will use the igraph package. We will also get rid of the self-loops.
## Warning: `graph.data.frame()` was deprecated in igraph 2.0.0.
## ℹ Please use `graph_from_data_frame()` instead.
## This warning is displayed once every 8 hours.
## Call `lifecycle::last_lifecycle_warnings()` to see where this warning was generated.
2.2 Network Structure
Now, we can start to explore the network structure. First of all, we can check the number of nodes and edges.
# Number of nodes
n_nodes <- igraph::vcount(network)
n_edges <- igraph::ecount(network)
print(paste("Number of nodes:", n_nodes))## [1] "Number of nodes: 563"
## [1] "Number of edges: 75120"
In network science, the number of nodes and edges is not enough to describe the network. We need to check the density of the network. The density of a network is the ratio between the number of edges and the maximum number of edges. The maximum number of edges is the number of nodes multiplied by the number of nodes minus one. It returns a value comprised between 0 and 1. The closer the density is to 1, the more dense the network is. The closer the density is to 0, the more sparse the network is.
# Density of the network
density <- igraph::edge_density(network)
print(paste("Density of the network:", density))## [1] "Density of the network: 0.237416483884629"
Overall, the network is not very dense. It means that the network overall is not very connected. However, we can check the density of the network for each party.
# Density of the network for each party
# density_party <- igraph::edge_density(network, membership = nodes$party)
# print(paste("Density of the network for each party:", density_party))Additionally, we can take a look at the diameter of the network. The diameter of a network is the maximum distance between two nodes. The distance can also be understood in social terms as the social distance (the number of steps) from one node to another. It is also called the geodesic distance.
## $vertices
## + 2/563 vertices, named, from 8e77d28:
## [1] secelainechao guamcongressman
##
## $distance
## [1] 6
In our case, the diameter of the network is 6 and the farthest vertices are those between Elaine Chao, who served as member of the Trump Cabinet as secretary of transportation and Michael San Nicolas, Democratic member of the House of Representatives from the Guam at-large district.