Homework 2

Descriptive Statistics and Reading and Transforming Data with R

Homework 2 questions based on Lectures 4 and 5.
Author

Byeong-Hak Choe

Published

September 30, 2026

Instructions

This homework covers Lectures 4 and 5. Answer all 35 questions: Questions 1–17 are conceptual questions (14 multiple-choice questions and three fill-in-the-blank questions), and Questions 18–35 ask you to fill in one R-code blank.

  • For each multiple-choice question, select the one best answer.
  • For each conceptual fill-in-the-blank question, enter only the requested term.
  • For each coding question, enter only the R code that replaces [?], not the full line or code block. Each question has one blank.
  • Run the data-loading code below once, then write and test your answers in an R script before submitting them.
  • Submit your answers through Brightspace. Check Brightspace for the due date.

For Questions 21–35, use this data frame. Questions 18–20 use the local-file scenario stated in Part III.

library(tidyverse)

# URL for the CSV file
CSV_url <- 'https://bcdanl.github.io/data/custdata_rev.csv'
custdata <- read_csv(CSV_url)

Variable descriptions

Each row represents one customer.

Variable Description
custid Unique customer identifier, stored as text.
sex Reported sex: Female or Male.
is_employed Employment status: TRUE = employed; FALSE = unemployed; NA = unknown or not applicable.
income Customer income in U.S. dollars.
marital_status Marital status, such as married, never married, divorced/separated, or widowed.
housing_type Housing situation, such as renting or owning a home.
recent_move Whether the customer moved in the past year: TRUE = yes; FALSE = no.
num_vehicles Number of vehicles in the household; 6 means six or more.
age Customer age in years.
state_of_res State of residence, including the District of Columbia.
gas_usage Monthly gas-bill field; values 4–999 represent dollar amounts, with special codes below.
health_ins Health-insurance status: TRUE = insured; FALSE = uninsured.

For gas_usage, 1 means included in rent or condo fees, 2 means included in the electricity payment, and 3 means no charge or gas not used. The gas questions use the recorded values. In any column, NA represents a missing or not-applicable value.

Reference: Customer data dictionary.

Part I: Multiple Choice

Question 1

Group A has ages 29, 30, and 31. Group B has ages 20, 30, and 40. Both groups have a mean age of 30. Which statement is correct?

  1. Group A has greater variability because its ages are closer together.
  2. Group B has greater variability because its ages are farther from the mean.
  3. The groups have equal variability because their means are equal.
  4. Variability cannot be compared unless the groups have different medians.

Question 2

The mean income in a customer dataset is higher this year than last year. What does this descriptive comparison establish by itself?

  1. A new marketing program caused the increase.
  2. Every customer’s income increased.
  3. The observed mean is higher, but the reason for the change is not established.
  4. The median income must also have increased.

Question 3

A customer’s gas_usage value appears beyond the upper fence of a boxplot. What is the most appropriate first interpretation?

  1. The value is certainly a data-entry error and must be deleted.
  2. The value is a possible outlier that should be checked in context.
  3. The value must be replaced by the mean gas usage.
  4. The value proves that all other gas-usage records are correct.

Question 4

For x <- c(4, 10, 18), what does R return from range(x)?

  1. c(4, 18), the minimum and maximum
  2. 14, the width from minimum to maximum
  3. 32, the sum of the values
  4. c(10, 14), the median and range width

Question 5

Suppose the third quartile (Q_3) of customer age is 62 years. Which interpretation is best?

  1. About 25% of ages are at or below 62.
  2. About 50% of ages are at or below 62.
  3. About 75% of ages are at or below 62.
  4. Every customer is older than 62.

Question 6

Ten customers each spend 20 dollars, and 90 customers each spend 50 dollars. Which expression gives the mean spending per customer across all 100 customers?

  1. (20+50)/2
  2. (10\times20+90\times50)/(10+90)
  3. (10+90)/(20+50)
  4. 20\times50

Question 7

Which description of a CSV file is correct?

  1. It is an R-only file that cannot be opened outside R.
  2. It is a plain-text table whose values are separated by commas; its first line often contains column names.
  3. It stores only one numeric vector and no column names.
  4. It automatically records the meaning of every variable and every missing value.

Question 8

If custdata is organized as tidy data with one row per customer, which description is correct?

  1. Each column is a customer, and each row is a variable.
  2. Each row is a customer, each column is a variable, and each cell holds one value.
  3. Each cell is a customer, and each row holds a complete variable.
  4. Each row must contain only numeric values.

Question 9

What does the native pipe |> do in custdata |> filter(age >= 65)?

  1. It assigns filter to the name custdata.
  2. It passes custdata as the first input to filter(), equivalent to filter(custdata, age >= 65).
  3. It joins custdata to a second data frame.
  4. It changes every age to 65.

Question 10

You want to keep only rows for customers aged 65 or older, then keep only the custid and age columns. Which functions perform these two tasks, in order?

  1. select(), then filter()
  2. filter(), then select()
  3. arrange(), then distinct()
  4. rename(), then arrange()

Question 11

A filter should keep customers who live in New York or New Jersey. Which operator combines the two state comparisons correctly?

  1. &, because each row must match both states
  2. |, because each row may match either state
  3. !, because each row must match neither state
  4. ==, because it combines two complete logical conditions

Question 12

Some values of is_employed are missing. Why should you use is.na(is_employed) rather than is_employed == NA to find them?

  1. NA is the same as FALSE.
  2. Comparing a value with NA using == returns NA, not a useful TRUE/FALSE missingness test.
  3. is.na() replaces missing values with zero.
  4. == NA works only when the column is numeric.

Question 13

How does arrange(custdata, state_of_res, desc(age)) order the rows?

  1. By age from oldest to youngest, then by state only when ages tie
  2. By state alphabetically; within each state, by age from oldest to youngest
  3. By state alphabetically; within each state, by age from youngest to oldest
  4. It removes rows with repeated states.

Question 14

What does distinct(custdata, state_of_res) return?

  1. One row for each unique state value, with a state_of_res column
  2. The number of customers in each state
  3. All original customer rows, sorted by state
  4. Only customers with missing state values

Part II: Conceptual Fill in the Blanks

For Questions 15–17, consider a new Posit Cloud project whose working directory has not been changed from the default, /cloud/project. A copy of custdata_rev.csv has been uploaded to a folder named data_files directly inside the project folder.

Question 15

R uses /cloud/project as the starting folder for interpreting relative pathnames. This starting folder is called R’s [Blank].

Enter only the two-word term that replaces the blank.

Question 16

The pathname /cloud/project/data_files/custdata_rev.csv gives the file’s full location, beginning at the root folder /. It is an [Blank] pathname.

Enter only the term that replaces the blank.

Question 17

With R’s working directory set to /cloud/project, the pathname data_files/custdata_rev.csv locates the file starting from that working directory. It is a [Blank] pathname.

Enter only the term that replaces the blank.

Part III: R Programming — Fill in the Code Blank

For Questions 18–20, use the same local-file scenario: R’s working directory is /cloud/project, and a copy of custdata_rev.csv is in the project’s data_files folder. Complete the code using these stated locations; you do not need to upload a file to answer these questions. Reading the CSV from the web in the Instructions does not save a local copy. For Questions 21–35, use the custdata data frame loaded from the web in the Instructions.

Question 18

Use an R function to obtain the current working directory. Enter only the function call that replaces [?].

current_folder <- [?]

Question 19

Read the local CSV using its absolute pathname. Enter only the pathname that replaces [?], including quotation marks.

custdata_absolute <- read_csv([?])

Question 20

Read the same local CSV using its relative pathname from /cloud/project. Enter only the pathname that replaces [?], including quotation marks.

custdata_relative <- read_csv([?])

Question 21

Count the number of rows in custdata. Enter only the code that replaces [?].

row_count <- [?]

Question 22

Extract the num_vehicles column as a vector. Enter only the code that replaces [?].

vehicle_counts <- [?]

Question 23

Calculate mean gas_usage, excluding missing values. Enter only the code that replaces [?].

mean_gas_usage <- [?]

Question 24

Calculate the third quartile (Q_3) of age using quantile(). Enter only the code that replaces [?].

third_quartile_age <- [?]

Question 25

Calculate the interquartile range of income. Enter only the code that replaces [?].

income_iqr <- [?]

Question 26

Calculate the sample standard deviation of income. Enter only the code that replaces [?].

income_sd <- [?]

Question 27

Keep rows for female customers aged 65 or older. Enter only the condition or conditions that replace [?].

older_women <- custdata |>
  filter([?])

Question 28

Keep rows for customers living in either New York or New Jersey. Enter only the condition that replaces [?].

ny_nj_customers <- custdata |>
  filter([?])

Question 29

Keep rows where is_employed is missing. Enter only the condition that replaces [?].

employment_missing <- custdata |>
  filter([?])

Question 30

Keep rows where gas_usage is not missing. Enter only the condition that replaces [?].

gas_usage_known <- custdata |>
  filter([?])

Question 31

Sort all customer rows by income from highest to lowest. Enter only the code that replaces [?].

highest_income_first <- custdata |>
  arrange([?])

Question 32

Sort by state_of_res alphabetically and, within each state, by age from oldest to youngest. Enter only the second sorting expression that replaces [?].

state_then_oldest <- custdata |>
  arrange(state_of_res, [?])

Question 33

Create a one-column data frame containing each distinct state of residence once. Enter only the column name that replaces [?].

states <- custdata |>
  distinct([?])

Question 34

Keep only the custid, age, and income columns, in that order. Enter only the column-name arguments that replace [?].

id_age_income <- custdata |>
  select([?])

Question 35

The pipeline below keeps female customers aged 65 or older, puts the highest incomes first, and keeps three columns. In the final step, rename income to annual_income. Enter only the function call that replaces [?].

older_women_income <- custdata |>
  filter(sex == "Female", age >= 65) |>
  arrange(desc(income)) |>
  select(custid, age, income) |>
  [?]
Back to top