/*! This file is auto-generated */ .wp-block-button__link{color:#fff;background-color:#32373c;border-radius:9999px;box-shadow:none;text-decoration:none;padding:calc(.667em + 2px) calc(1.333em + 2px);font-size:1.125em}.wp-block-file__button{background:#32373c;color:#fff;text-decoration:none} Problem 68 Making Friends Online A survey c... [FREE SOLUTION] | 91Ó°ÊÓ

91Ó°ÊÓ

Making Friends Online A survey conducted in March 2015 asked 1060 teens to estimate, on average, the number of friends they had made online. While \(43 \%\) had not made any friends online, a small number of the teens had made many friends online. (a) Do you expect the distribution of number of friends made online to be symmetric, skewed to the right, or skewed to the left? (b) Two measures of center for this distribution are 1 friend and 5.3 friends. \({ }^{31}\) Which is most likely to be the mean and which is most likely to be the median? Explain your reasoning.

Short Answer

Expert verified
(a) The distribution of number of friends made online is likely skewed to the right. (b) 5.3 friends is the mean and 1 friend is the median.

Step by step solution

01

Understanding Distribution Skewness

Skewness is a measure of asymmetricity in a data set. In symmetric distribution, values are evenly distributed around the mean. Now, considering the problem, a large percentage (43%) of teens did not make any friends online, which places a lot of responses at the extreme left of the distribution and a smaller percentage made many friends, which represents fewer values to the extreme right of the distribution. This suggests a skew to the right.
02

Identify Mean and Median

A skew to the right means that there are a few very high values that 'pull' the average upwards. In skewed distribution (either right or left), mean is influenced by outliers while median is not. Therefore, mean is typically larger than median in a right-skewed distribution.
03

Applying Mean and Median to Problem

In this case, given the measures of center (1 friend and 5.3 friends), it's reasonable to consider the higher value, 5.3 friends, as the mean (affected by high outliers: teens who made many friends) and the lesser value, 1 friend, as the median (representing the middle value, unaffected by outliers).

Unlock Step-by-Step Solutions & Ace Your Exams!

  • Full Textbook Solutions

    Get detailed explanations and key concepts

  • Unlimited Al creation

    Al flashcards, explanations, exams and more...

  • Ads-free access

    To over 500 millions flashcards

  • Money-back guarantee

    We refund you if you fail your exam.

Over 30 million students worldwide already upgrade their learning with 91Ó°ÊÓ!

Key Concepts

These are the key concepts you need to understand to accurately answer the question.

Mean vs Median
When we look at a dataset, we often use numbers like the mean and median to describe the "center" or "average" of that data.
Both the mean and median help to summarize the data, but they do it in slightly different ways.

The **mean** is what many people think of as the "average." It's calculated by adding up all the numbers in a set and then dividing by how many numbers there are. However, the mean can be greatly influenced by extremely high or low numbers, known as outliers.
  • This susceptibility to outliers might make the mean not truly representative of the dataset.
On the other hand, the **median** is the middle number in a sorted list of numbers.
It's the point where half the numbers are higher and half are lower.
  • The median isn't affected by outliers because it relies purely on position in the data, making it a robust measure, especially in skewed distributions.
In a right-skewed distribution, where there are a few very high values, the mean gets "pulled" to the right by those large values, ending up higher than the median.
Data Asymmetry
Data asymmetry refers to how data points in a dataset are not evenly distributed on either side of a central point like the mean.
In symmetric distributions, values are spread evenly, looking the same on the left and right of the center.

However, in many real-world datasets, perfect symmetry is rare.
  • Right-skewed distributions have a "tail" that extends more widely to the right.
  • Left-skewed distributions extend more widely to the left.
Understanding whether data is symmetric or asymmetric helps us decide on what summary statistics to use.
For example, in the case of the teen friendship survey, the 43% of teens who had zero online friends heavily influences the shape of the distribution.
This causes the values to cluster heavily on the left while the right "tail" stretches due to teens who made many online friends, indicating data asymmetry.
Right Skewness
Right skewness, also referred to as positive skewness, occurs when the majority of data points cluster to the left, with the tail extending to the right.

In the context of the teen survey about online friends, 43% of respondents reported making no online friends, resulting in many values at zero or low numbers. However, since some teens reported a high number of online friends, these few higher values create a long tail on the right.
  • This right-skewed nature suggests looking closer at data points on the extreme high side for insights.
Right-skewed distributions have some important characteristics:
  • They often result in the mean being higher than the median because of the impact of the "tail" that's pulling the mean upwards.
  • They can indicate that there are outliers or uncommon events worth investigating separately.
Understanding skewness, particularly right skewness, can signal when extra caution is needed in interpreting data averages like the mean.

One App. One Place for Learning.

All the tools & learning materials you need for study success - in one app.

Get started for free

Most popular questions from this chapter

Use the \(95 \%\) rule and the fact that the summary statistics come from a distribution that is symmetric and bell-shaped to find an interval that is expected to contain about \(95 \%\) of the data values. A bell-shaped distribution with mean 200 and standard deviation 25.

In Exercise 2.120 on page \(92,\) we discuss a study in which the Nielsen Company measured connection speeds on home computers in nine different countries in order to determine whether connection speed affects the amount of time consumers spend online. \(^{69}\) Table 2.29 shows the percent of Internet users with a "fast" connection (defined as \(2 \mathrm{Mb}\) or faster) and the average amount of time spent online, defined as total hours connected to the Web from a home computer during the month of February 2011. The data are also available in the dataset GlobalInternet. (a) What would a positive association mean between these two variables? Explain why a positive relationship might make sense in this context. (b) What would a negative association mean between these two variables? Explain why a negative relationship might make sense in this context. $$ \begin{array}{lcc} \hline \text { Country } & \begin{array}{c} \text { Percent Fast } \\ \text { Connection } \end{array} & \begin{array}{l} \text { Hours } \\ \text { Online } \end{array} \\ \hline \text { Switzerland } & 88 & 20.18 \\ \text { United States } & 70 & 26.26 \\ \text { Germany } & 72 & 28.04 \\ \text { Australia } & 64 & 23.02 \\ \text { United Kingdom } & 75 & 28.48 \\ \text { France } & 70 & 27.49 \\ \text { Spain } & 69 & 26.97 \\ \text { Italy } & 64 & 23.59 \\ \text { Brazil } & 21 & 31.58 \\ \hline \end{array} $$ (c) Make a scatterplot of the data, using connection speed as the explanatory variable and time online as the response variable. Is there a positive or negative relationship? Are there any outliers? If so, indicate the country associated with each outlier and describe the characteristics that make it an outlier for the scatterplot. (d) If we eliminate any outliers from the scatterplot, does it appear that the remaining countries have a positive or negative relationship between these two variables? (e) Use technology to compute the correlation. Is the correlation affected by the outliers? (f) Can we conclude that a faster connection speed causes people to spend more time online?

When honeybee scouts find a food source or a nice site for a new home, they communicate the location to the rest of the swarm by doing a "waggle dance." 74 They point in the direction of the site and dance longer for sites farther away. The rest of the bees use the duration of the dance to predict distance to the site. Table 2.32 Duration of \(a\) honeybee waggle dance to indicate distance to the source $$\begin{array}{cc} \hline \text { Distance } & \text { Duration } \\ \hline 200 & 0.40 \\\250 & 0.45 \\ 500 & 0.95 \\\950 & 1.30 \\ 1950 & 2.00 \\\3500 & 3.10 \\\4300 & 4.10 \\\\\hline\end{array}$$ Table 2.32 shows the distance, in meters, and the duration of the dance, in seconds, for seven honeybee scouts. \(^{75}\) This information is also given in HoneybeeWaggle. (a) Which is the explanatory variable? Which is the response variable? (b) Figure 2.70 shows a scatterplot of the data. Does there appear to be a linear trend in the data? If so, is it positive or negative? (c) Use technology to find the correlation between the two variables. (d) Use technology to find the regression line to predict distance from duration. (e) Interpret the slope of the line in context. (f) Predict the distance to the site if a honeybee does a waggle dance lasting 1 second. Lasting 3 seconds.

Fiber in the Diet The number of grams of fiber eaten in one day for a sample of ten people are \(\begin{array}{ll}10 & 11\end{array}\) \(\begin{array}{ll}11 & 14\end{array}\) \(\begin{array}{llllll}15 & 17 & 21 & 24 & 28 & 115\end{array}\) (a) Find the mean and the median for these data. (b) The value of 115 appears to be an obvious outlier. Compute the mean and the median for the nine numbers with the outlier excluded. (c) Comment on the effect of the outlier on the mean and on the median.

Ages of Husbands and Wives Suppose we record the husband's age and the wife's age for many randomly selected couples. (a) What would it mean about ages of couples if these two variables had a negative relationship? (b) What would it mean about ages of couples if these two variables had a positive relationship? (c) Which do you think is more likely, a negative or a positive relationship? (d) Do you expect a strong or a weak relationship in the data? Why? (e) Would a strong correlation imply there is an association between husband age and wife age?

See all solutions

Recommended explanations on Math Textbooks

View all explanations

What do you think about this solution?

We value your feedback to improve our textbook solutions.

Study anywhere. Anytime. Across all devices.