library(tidyverse)
# URL for the CSV file
CSV_url <- 'https://bcdanl.github.io/data/custdata_rev.csv'
custdata <- read_csv(CSV_url)Homework 2
Descriptive Statistics and Reading and Transforming Data with R
Instructions
This homework covers Lectures 4 and 5. Answer all 35 questions: Questions 1–17 are conceptual questions (14 multiple-choice questions and three fill-in-the-blank questions), and Questions 18–35 ask you to fill in one R-code blank.
- For each multiple-choice question, select the one best answer.
- For each conceptual fill-in-the-blank question, enter only the requested term.
- For each coding question, enter only the R code that replaces
[?], not the full line or code block. Each question has one blank. - Run the data-loading code below once, then write and test your answers in an R script before submitting them.
- Submit your answers through Brightspace. Check Brightspace for the due date.
For Questions 21–35, use this data frame. Questions 18–20 use the local-file scenario stated in Part III.
Variable descriptions
Each row represents one customer.
| Variable | Description |
|---|---|
custid |
Unique customer identifier, stored as text. |
sex |
Reported sex: Female or Male. |
is_employed |
Employment status: TRUE = employed; FALSE = unemployed; NA = unknown or not applicable. |
income |
Customer income in U.S. dollars. |
marital_status |
Marital status, such as married, never married, divorced/separated, or widowed. |
housing_type |
Housing situation, such as renting or owning a home. |
recent_move |
Whether the customer moved in the past year: TRUE = yes; FALSE = no. |
num_vehicles |
Number of vehicles in the household; 6 means six or more. |
age |
Customer age in years. |
state_of_res |
State of residence, including the District of Columbia. |
gas_usage |
Monthly gas-bill field; values 4–999 represent dollar amounts, with special codes below. |
health_ins |
Health-insurance status: TRUE = insured; FALSE = uninsured. |
For gas_usage, 1 means included in rent or condo fees, 2 means included in the electricity payment, and 3 means no charge or gas not used. The gas questions use the recorded values. In any column, NA represents a missing or not-applicable value.
Reference: Customer data dictionary.
Part I: Multiple Choice
Question 1
Group A has ages 29, 30, and 31. Group B has ages 20, 30, and 40. Both groups have a mean age of 30. Which statement is correct?
- Group A has greater variability because its ages are closer together.
- Group B has greater variability because its ages are farther from the mean.
- The groups have equal variability because their means are equal.
- Variability cannot be compared unless the groups have different medians.
Question 2
The mean income in a customer dataset is higher this year than last year. What does this descriptive comparison establish by itself?
- A new marketing program caused the increase.
- Every customer’s income increased.
- The observed mean is higher, but the reason for the change is not established.
- The median income must also have increased.
Question 3
A customer’s gas_usage value appears beyond the upper fence of a boxplot. What is the most appropriate first interpretation?
- The value is certainly a data-entry error and must be deleted.
- The value is a possible outlier that should be checked in context.
- The value must be replaced by the mean gas usage.
- The value proves that all other gas-usage records are correct.
Question 4
For x <- c(4, 10, 18), what does R return from range(x)?
c(4, 18), the minimum and maximum14, the width from minimum to maximum32, the sum of the valuesc(10, 14), the median and range width
Question 5
Suppose the third quartile (Q_3) of customer age is 62 years. Which interpretation is best?
- About 25% of ages are at or below 62.
- About 50% of ages are at or below 62.
- About 75% of ages are at or below 62.
- Every customer is older than 62.
Question 6
Ten customers each spend 20 dollars, and 90 customers each spend 50 dollars. Which expression gives the mean spending per customer across all 100 customers?
- (20+50)/2
- (10\times20+90\times50)/(10+90)
- (10+90)/(20+50)
- 20\times50
Question 7
Which description of a CSV file is correct?
- It is an R-only file that cannot be opened outside R.
- It is a plain-text table whose values are separated by commas; its first line often contains column names.
- It stores only one numeric vector and no column names.
- It automatically records the meaning of every variable and every missing value.
Question 8
If custdata is organized as tidy data with one row per customer, which description is correct?
- Each column is a customer, and each row is a variable.
- Each row is a customer, each column is a variable, and each cell holds one value.
- Each cell is a customer, and each row holds a complete variable.
- Each row must contain only numeric values.
Question 9
What does the native pipe |> do in custdata |> filter(age >= 65)?
- It assigns
filterto the namecustdata. - It passes
custdataas the first input tofilter(), equivalent tofilter(custdata, age >= 65). - It joins
custdatato a second data frame. - It changes every age to 65.
Question 10
You want to keep only rows for customers aged 65 or older, then keep only the custid and age columns. Which functions perform these two tasks, in order?
select(), thenfilter()filter(), thenselect()arrange(), thendistinct()rename(), thenarrange()
Question 11
A filter should keep customers who live in New York or New Jersey. Which operator combines the two state comparisons correctly?
&, because each row must match both states|, because each row may match either state!, because each row must match neither state==, because it combines two complete logical conditions
Question 12
Some values of is_employed are missing. Why should you use is.na(is_employed) rather than is_employed == NA to find them?
NAis the same asFALSE.- Comparing a value with
NAusing==returnsNA, not a usefulTRUE/FALSEmissingness test. is.na()replaces missing values with zero.== NAworks only when the column is numeric.
Question 13
How does arrange(custdata, state_of_res, desc(age)) order the rows?
- By age from oldest to youngest, then by state only when ages tie
- By state alphabetically; within each state, by age from oldest to youngest
- By state alphabetically; within each state, by age from youngest to oldest
- It removes rows with repeated states.
Question 14
What does distinct(custdata, state_of_res) return?
- One row for each unique state value, with a
state_of_rescolumn - The number of customers in each state
- All original customer rows, sorted by state
- Only customers with missing state values
Part II: Conceptual Fill in the Blanks
For Questions 15–17, consider a new Posit Cloud project whose working directory has not been changed from the default, /cloud/project. A copy of custdata_rev.csv has been uploaded to a folder named data_files directly inside the project folder.
Question 15
R uses /cloud/project as the starting folder for interpreting relative pathnames. This starting folder is called R’s [Blank].
Enter only the two-word term that replaces the blank.
Question 16
The pathname /cloud/project/data_files/custdata_rev.csv gives the file’s full location, beginning at the root folder /. It is an [Blank] pathname.
Enter only the term that replaces the blank.
Question 17
With R’s working directory set to /cloud/project, the pathname data_files/custdata_rev.csv locates the file starting from that working directory. It is a [Blank] pathname.
Enter only the term that replaces the blank.
Part III: R Programming — Fill in the Code Blank
For Questions 18–20, use the same local-file scenario: R’s working directory is /cloud/project, and a copy of custdata_rev.csv is in the project’s data_files folder. Complete the code using these stated locations; you do not need to upload a file to answer these questions. Reading the CSV from the web in the Instructions does not save a local copy. For Questions 21–35, use the custdata data frame loaded from the web in the Instructions.
Question 18
Use an R function to obtain the current working directory. Enter only the function call that replaces [?].
current_folder <- [?]Question 19
Read the local CSV using its absolute pathname. Enter only the pathname that replaces [?], including quotation marks.
custdata_absolute <- read_csv([?])Question 20
Read the same local CSV using its relative pathname from /cloud/project. Enter only the pathname that replaces [?], including quotation marks.
custdata_relative <- read_csv([?])Question 21
Count the number of rows in custdata. Enter only the code that replaces [?].
row_count <- [?]Question 22
Extract the num_vehicles column as a vector. Enter only the code that replaces [?].
vehicle_counts <- [?]Question 23
Calculate mean gas_usage, excluding missing values. Enter only the code that replaces [?].
mean_gas_usage <- [?]Question 24
Calculate the third quartile (Q_3) of age using quantile(). Enter only the code that replaces [?].
third_quartile_age <- [?]Question 25
Calculate the interquartile range of income. Enter only the code that replaces [?].
income_iqr <- [?]Question 26
Calculate the sample standard deviation of income. Enter only the code that replaces [?].
income_sd <- [?]Question 27
Keep rows for female customers aged 65 or older. Enter only the condition or conditions that replace [?].
older_women <- custdata |>
filter([?])Question 28
Keep rows for customers living in either New York or New Jersey. Enter only the condition that replaces [?].
ny_nj_customers <- custdata |>
filter([?])Question 29
Keep rows where is_employed is missing. Enter only the condition that replaces [?].
employment_missing <- custdata |>
filter([?])Question 30
Keep rows where gas_usage is not missing. Enter only the condition that replaces [?].
gas_usage_known <- custdata |>
filter([?])Question 31
Sort all customer rows by income from highest to lowest. Enter only the code that replaces [?].
highest_income_first <- custdata |>
arrange([?])Question 32
Sort by state_of_res alphabetically and, within each state, by age from oldest to youngest. Enter only the second sorting expression that replaces [?].
state_then_oldest <- custdata |>
arrange(state_of_res, [?])Question 33
Create a one-column data frame containing each distinct state of residence once. Enter only the column name that replaces [?].
states <- custdata |>
distinct([?])Question 34
Keep only the custid, age, and income columns, in that order. Enter only the column-name arguments that replace [?].
id_age_income <- custdata |>
select([?])Question 35
The pipeline below keeps female customers aged 65 or older, puts the highest incomes first, and keeps three columns. In the final step, rename income to annual_income. Enter only the function call that replaces [?].
older_women_income <- custdata |>
filter(sex == "Female", age >= 65) |>
arrange(desc(income)) |>
select(custid, age, income) |>
[?]