This page is part of my personal academic record, not an official course website.

Course description

This course introduces the application of supercomputing to statistical data analysis, particularly on big data. Implementations of various statistical methodologies within a parallel-computing framework are demonstrated through all lectures. The course will cover (1) parallel computing basics, including architectures of interconnection networks, communication methodologies, algorithm and performance measurements, and (2) their applications to modern data mining techniques, including modern variable selection/dimension reduction, linear/logistic regression, tree-based classification methods, kernel-based methods, nonlinear statistical models, and model inference/resampling methods.

Textbooks and resources

Supplementary textbooks

  • Applied Parallel Computing by Yuefan Deng, 2012, World Scientific
  • The Elements of Statistical Learning: Data Mining, Inference, and Prediction by Trevor Hastie, Robert Tibshirani and Jerome Friedman, 2nd edition, 2016, Springer
  • Mining of Massive Datasets by Jure Leskovec, Anand Rajaraman and Jeffrey David Ullman, 3rd edition, Cambridge University Press

Learning outcomes

  1. Demonstrate knowledge of parallel computing basics:
    • Node architecture, central processing units, and accelerators;
    • Distributed- and shared-memory computing.
  2. Demonstrate skills with software architecture and R:
    • Communication patterns and protocols;
    • Process creation and management;
    • MapReduce framework;
    • Hadoop in R.
  3. Demonstrate mastery of basic tools for big data analysis:
    • Linear regression;
    • Logistic regression;
    • Dimension reduction.
  4. Demonstrate understanding of advanced methods for big data analysis:
    • Classification and regression trees;
    • Random forest;
    • Gradient boosting;
    • Support vector machine;
    • Neural network.
  5. Demonstrate understanding of model selection and performance evaluation:
    • Best subset; forward selection; backward selection;
    • Cross-validation;
    • Bootstrap.

Back to Graduate Studies