Data Types, Databases, and Big Data
October 7, 2026
Discuss: Which source would help you count purchases? Which could help explain why customers liked a product?
Defined rows and columns.
Example: an order table with an ID, date, and total.
Labels or tags organize information, but records may differ.
Example: an app's labeled data response.
The main content does not follow a fixed table layout.
Example: review text, photos, or videos.
| Kind of variable | What the values mean | Useful comparison |
|---|---|---|
| Nominal | Categories without a ranking | Same or different |
| Ordinal | Categories with a ranking | Higher or lower |
| Interval | Equal differences; no meaningful zero | How much higher or lower |
| Ratio | Equal differences and a meaningful zero | How much; how many times as much |
| Customer | Channel |
|---|---|
| C01 | Website |
| C02 | Store |
| C03 | App |
Discuss: If we label these channels 1, 2, and 3, does that make “average channel = 2” meaningful?
| Response | Order |
|---|---|
| Dissatisfied | Lower |
| Neutral | Middle |
| Satisfied | Higher |
Discuss: On a 1–5 satisfaction scale, does 4 mean “twice as satisfied” as 2?
| Day | Temperature (°F) |
|---|---|
| Monday | 70 |
| Tuesday | 80 |
| Wednesday | 90 |
Check: Which comparison works here—“10°F warmer” or “twice as hot”?
| Customer | Minutes |
|---|---|
| C01 | 0 |
| C02 | 15 |
| C03 | 30 |
| Variable | Example | How to interpret it |
|---|---|---|
| Customer ID | 101, 102, 103 | Labels identifying people |
| Satisfaction score | 1–5 | Ordered responses; gaps may differ |
| Number of orders | 0, 1, 2, 3 | A count with a meaningful zero |
Discuss: Why would an average customer ID be unhelpful, even if R can calculate it?
Note
SQL (Structured Query Language) is commonly used to query database tables. In this course, we use R to practice working with tables.
Reference: IBM · Relational databases
| Item | Role | Example in our workflow |
|---|---|---|
| CSV file | Saves one table as plain text | A downloaded sales export |
| Database | Stores and organizes data for ongoing use | A retailer’s customer and order tables |
| R data frame | Holds a table in an R session | Data read with read_csv() |
| Column | Meaning | Example rule |
|---|---|---|
order_id |
Identifies an order | Required and unique |
customer_id |
Identifies the customer | Uses the agreed ID format |
total |
Order amount in dollars | Numeric; unit documented |
Discuss: A total of $4,000 is numeric. What else would you check before trusting it?
orders: one row per order| order_id | customer_id | total |
|---|---|---|
| 101 | C01 | 25 |
| 102 | C02 | 40 |
| 103 | C01 | 15 |
| 104 | C03 | 30 |
customers: one row per customer| customer_id | city |
|---|---|
| C01 | Rochester |
| C02 | Buffalo |
| C04 | Albany |
Discuss: Which column could connect an order to its customer’s city?
customer_id.orderscustomer_id tells us which customer placed each order.
C01 appears in two orders because the same customer ordered twice.
customerscustomer_id identifies the customer whose city is recorded.
C01 matches the customer record for Rochester.
orders is our main table; we want to keep its orders.customers supplies the city to attach to each matching order.orders, the table on the left.by = "customer_id" tells R which column’s values to match.left_join() keeps every order and adds city from matching customer records.NA.Reference: dplyr · Mutating joins
| order_id | customer_id | total | city |
|---|---|---|---|
| 101 | C01 | 25 | Rochester |
| 102 | C02 | 40 | Buffalo |
| 103 | C01 | 15 | Rochester |
| 104 | C03 | 30 | NA |
NA because C03 has no customer match.Check: Why does C04’s city, Albany, not appear in this result?
customers accidentally lists two cities for C01.| customer_id | city |
|---|---|
| C01 | Rochester |
| C01 | Syracuse |
| order_id | customer_id | city |
|---|---|---|
| 101 | C01 | Rochester |
| 101 | C01 | Syracuse |
Discuss: Which city belongs to C01? What evidence would you check?
Predict: If we start with customers instead, which records would we be trying to keep?
Copy data from its sources.
Example: collect order exports and customer records.
Clean and connect the data.
Example: standardize IDs and attach customer cities.
Save prepared data to a destination.
Example: store a reporting table in a warehouse.
Reference: IBM · ETL
read_csv() is one way to import a source table.Discuss: Your file has no orders for Sunday. How could you tell whether the store was closed or the export is incomplete?
| Issue in the source | What to check or change |
|---|---|
" C01 " versus "C01" |
Remove unintended spaces in IDs |
10/07/2026 |
Confirm the date convention before converting it |
| Amounts from different countries | Check currencies before combining totals |
| Repeated or missing customer IDs | Investigate before joining |
filter() keeps chosen rows, select() selects columns, and left_join() connects tables.Prepared data is saved to a destination, such as a database or data warehouse.
Other people can retrieve it for reporting and analysis.
We hold a prepared table in an R data frame.
Saving it as a CSV is a simple way to retain a result outside the R session.
| V | Question to ask |
|---|---|
| Volume | How much data must we store or process? |
| Velocity | How quickly does new data arrive or need to be used? |
| Variety | What different formats and sources must we combine? |
| Veracity | How reliable and complete is the data? |
| Value | What useful decision can this data support? |
Reference: IBM · Big data
| Unit | Decimal size in bytes |
|---|---|
| Kilobyte (kB) | 1,000 |
| Megabyte (MB) | 1,000,000 |
| Gigabyte (GB) | 1,000,000,000 |
| Terabyte (TB) | 1,000,000,000,000 |
| Petabyte (PB) | 1,000,000,000,000,000 |
New orders and driver locations arrive continuously.
A dispatcher needs recent information to assign a nearby driver.
Managers review sales across the completed month.
They need consistent totals, but not necessarily an update every second.
Discuss: Could yesterday’s accurate driver locations help dispatch an order right now?
Order IDs, dates, quantities, and prices.
Written comments explaining a shopping experience.
An image showing a product or possible defect.
Discuss: If every value fits the expected data type, which of these problems could still remain?
Discuss: What could you miss if you chose next month’s stock using sales totals alone?
| Need | Possible response |
|---|---|
| More data than one computer can hold | Use larger storage or spread it across computers |
| A task takes too long | Divide suitable work across several computers |
| Many people need reliable access | Use a shared database with access rules |
Prepared, connected data with shared definitions.
Reference: IBM · Data warehouses
| Operational database | Data warehouse | |
|---|---|---|
| Main purpose | Support daily operations | Support analysis and reporting |
| Typical task | Record an order or update stock | Compare sales across stores and months |
| Data focus | Records needed by the application | Prepared data combined across sources |
Illustrative image
1992 milestone: Walmart’s Teradata data warehouse became the first to reach 1 TB.
Possible analytics use: Combine entry/exit counts with checkout activity by store and hour to identify busy periods and plan checkout staffing.
Discuss: Which data would help suggest a product? Which collection raises greater privacy concerns?