In a graph database, a _______ is a data entity represented by a node.
- Document
- Edge
- Relationship
- Vertex
In a graph database, a "Vertex" is a data entity represented by a node. A vertex typically contains properties that describe the entity, and the relationships between vertices define the connections in the graph.
A manufacturing company wants to calculate the average production output per factory location. Which data modeling technique would you recommend for this scenario?
- Entity-Relationship Diagram
- Fact and Dimension Tables
- Snowflake Schema
- Star Schema
To calculate the average production output per factory location, the recommended data modeling technique is to use Fact and Dimension Tables. This approach involves creating a fact table containing production data and dimension tables providing details about factory locations, enabling efficient analysis.
What are clustering techniques used for in relational schema design?
- Creating composite keys
- Grouping related tables together on disk
- Implementing referential integrity
- Reducing data redundancy
Clustering techniques in relational schema design involve grouping related tables together on disk. This can enhance query performance by minimizing disk I/O when retrieving data from interconnected tables in a query.
A _______ constraint is used to ensure that a column value meets specific criteria.
- Check
- Foreign
- Primary
- Unique
Detailed A check constraint is used to ensure that a column value meets specific criteria or conditions. This helps in maintaining data accuracy and consistency by defining rules that must be satisfied for data in a column.
A random variable that takes a finite or countably infinite number of values is known as a ________ random variable.
- Continuous
- Dependent
- Discrete
- Normal
A discrete random variable is one which may take on only a countable number of distinct values and thus can be quantified. For example, you can count the change in your pocket. You can count the money in your bank account. You can count the number of heads in 50 coin tosses. These are all examples of discrete random variables.
A situation where two or more independent variables in a regression model are highly correlated is known as ________.
- autocorrelation
- heteroscedasticity
- homoscedasticity
- multicollinearity
Multicollinearity refers to a situation in which two or more independent variables in a regression model are highly linearly related. This can lead to unstable estimates of the regression coefficients and make it difficult to assess the effect of independent variables on the dependent variable.
How does multiple linear regression differ from simple linear regression?
- Multiple linear regression cannot handle categorical variables, simple linear regression can
- Multiple linear regression is not suitable for prediction tasks
- Multiple linear regression requires a larger dataset
- Multiple linear regression uses multiple independent variables, simple linear regression only uses one
The main difference between simple and multiple linear regression is the number of independent variables. While simple linear regression uses only one independent variable to predict the dependent variable, multiple linear regression uses two or more independent variables to predict the dependent variable.
What does the residual plot tell you in a simple linear regression analysis?
- It shows the distribution of residuals and can help identify non-linearity, unequal error variances, and outliers
- It shows the distribution of the independent variable
- It shows the relationship between the dependent and independent variables
- It tells you the strength of the correlation
A residual plot is a graph that shows the residuals on the vertical axis and the independent variable on the horizontal axis. It helps to identify non-linearity, unequal error variances (heteroscedasticity), and outliers. If the points in a residual plot are randomly dispersed around the horizontal axis, a linear regression model is appropriate for the data; otherwise, a non-linear model is more appropriate.
Two events are said to be ________ if the occurrence of one does not affect the probability of the occurrence of the other.
- Dependent
- Exhaustive
- Independent
- Mutually exclusive
Two events are said to be "independent" if the occurrence of one does not affect the probability of the occurrence of the other. For example, if you toss a coin twice, the outcome of the first toss doesn't affect the outcome of the second toss, so the two events are independent.
How does sample size impact the Mann-Whitney U test?
- Larger sample sizes make the test less reliable
- Larger sample sizes make the test more reliable
- Only equal sample sizes can be used in the test
- Sample size has no impact on the test
Larger sample sizes make the Mann-Whitney U test more reliable. As with most statistical tests, a larger sample size increases the power of the test, which is the probability that it will correctly reject a false null hypothesis.