2.Weight loss: In a study to determine whether counseling could help people lose weight, a sample of people experienced a
eople experienced a group-based behavioral intervention, which involved weekly meetings with a trained interventionist for a period of six months. The following data are the numbers of pounds lost for 14 people, based on means and standard deviations given in the article. Assume the population is approximately normal. Perform a hypothesis test to determine whether the mean weight loss is greater than 20 pounds. Use the =α0.10 level of significance and the critical value method.
22.5 28.5 7.6 24.1 21.5 12.9 17.3
21.2 37.6 33.8 12.1 36.3 24.1 19.4
You are asked to carry out a study on behalf of a business analytics specialised consultancy on a subsample
on a subsample of weekly data from Randall’s Supermarket, one of the biggest in the UK. Randall’s marketing management team wishes to identify trends and patterns in a sample of weekly data collected for a number of their loyalty cardholders during a 26-week period. The data includes information on the customers’ gender, age, shopping frequency per week and shopping basket price. Randall’s operates two different types of stores (convenient stores and superstores) but they also sell to customers via an online shopping platform. The collected data are from all three different types of stores. Finally, the data provides information on the consistency of the customer’s shopping basket regarding the type of products purchased. These can vary from value products, to brand as well as the supermarket’s own high-quality product series Randall’s Top. As a business analyst you are required to analyse those data, make any necessary modifications in order to determine whether for any single customer it is possible to predict the value of their shopping basket.
Randall’s marketing management team is only interested in identifying whether the spending of the potential customer will fall in one of three possible groups including:
• Low spender (shopping basket value of £25 or less)
• Medium Spender (shopping basket value between £25.01 and £70) and
• High spenders (shopping basket greater than £70)
For the purpose of your analysis you are provided with the data set Randall’s.xls. You have to decide, which method is appropriate to apply for the problem under consideration and undertake the necessary analysis. Once you have completed this analysis, write a report for the Randall’s marketing management team summarising your findings but also describing all necessary steps undertaken in the analysis. The manager is a competent business analyst himself/herself so the report can include technical terms, although you should not exceed five pages. Screenshots and supporting materials can be included in the appendix.
After completing your analysis, you should submit a report that consists of two parts. Part A being a non-technical summary of your findings and Part B a detailed report of the analysis undertaken with more details.
Part A: A short report for the Head of Randall’s Marketing Management (20 per cent). This should briefly explain the aim of the project, a clear summary and justification of the methods considered as well as an overview of the results.
Although, the Head of Randall’s Marketing Management team who will receive this summary is a competent business analytics practitioner, the majority of the other team members have little knowledge of statistical modelling and want to know nothing about the technical and statistical underpinning of the techniques used in this analysis. This report should be no more than two sides of A4 including graphs, tables, etc. In this report you should include all the objectives of this analysis, summary of data and results as well as your recommendations (if any).
Part B: A technical report on the various stages of the analysis (80 per cent).
The analysis should be carried out using the range of analytics tools discussed:
• SPSS Statistics
Ensure that the exercise references:
• Binary and multinomial logistic regression
• Linear vs Logistic regression
• Logit Model with odds Ratio
• Co-efficients and Chi Squared
• MLR co-efficients
• Assessing usefulness of MLR model
• Interpreting a model
• Assessing over-all model fit with Psuedo R-Squared measures
• Classification accuracy (Hit Ratio)
• Wald Statistic
• Odd ratio exp(B)
• Ratio of the probability of an event happening vs not happening
• Ratio of the odds after a unit change in the predictor to the original odds
• Residuals analysis
• Cook’s distance
• Adequacy (with variance inflation factor VIF and tolerance statistic)
• Outliers and influential points cannot just be removed. We need to check them (typo? – unusual data?)
• Check for multicollinearity
Write a short and concise report to explain the technical detail of what you have done for each step of the analysis.
The report should also cover the following information:
• Any type of analysis that might be useful and check whether the main assumptions behind the analyses do not hold or cannot be
• Give evidence of the understanding of the statistical tools that you are using. For example, comment on the model selection procedure and the coefficient interpretation, e.g. comment on the interpretation of the logistic regression coefficients if such a method is used and provide an example of
• Conclusions and explanation, in non-technical terms, of the main points
4.4This week, you'll apply your knowledge of data collections for sentiment analysis, a common technique applied to movie, product, and
ue applied to movie, product, and business reviews, as well as social media posts.
For this assignment, you'll determine whether teacher reviews are positive or negative, using real reviews from Rate my Professor.
Modify this wordfreq.py (Links to an external site.)Links to an external site. program to evaluate whether a particular review is positive, negative, or neutral.
You can use these collections of positivePreview the document and negativePreview the document words (in files folder of Canvas). Note the files contain non-ascii characters that you'll need to accommodate like so:
negWords = open('negative-words.txt','r', encoding='utf-8', errors='ignore').read().splitlines()[35:]
Your modified program should:
exclude 'stop' words from your word counts, using the the below list;
a, an, and, as, at, be, but, etc, for, in, it, its, is, of, or, so, such, the, this, to, with
print the remaining top 25 words, along with their frequency,
print the top 5 positive and top 5 negative words, along with their frequency,
calculate and display a sentiment score for the teacher, where the score is incremented (+1) for each positive word in the review and decremented (-1) for each negative word
- Use this set of real-life teacher reviews