Fall 2024 Reading List

The written exam of the DAIS Qual Exam in Fall 2024 will be held on Friday,  Oct. 4, 2024, at 3pm-7pm in room 1218 Siebel Center, which is in the basement of Siebel Center (floor map: https://facilityaccessmaps.fs.illinois.edu/archibus/schema/ab-products/essential/workplace/?blId=0563&flId=00). 

This reading list consists of  multiple topic sections, each containing 2-3 papers.  The questions in the written exam will be based on the papers listed here, with 1-2 questions related to each section. If a section has two papers, you can usually expect to see one question related to the section in the qual exam, while if a section has three papers, you can usually expect to see two questions related to the section. You only need to answer four of those questions in the exam, so there is no need for you to read every paper. Instead, it would make sense for you to browse through the list and identify up to 4 sections that have papers that you are most familiar with or most comfortable with reading, and then focus on reading/digesting those papers. In general, you will likely find some sections to be closer to your interests or background than others, and you can focus more on reading the papers in those a few sections that seem to be closest to your research interests.    

Section 1

  • Mengting Wan, Tara Safavi, Sujay Kumar Jauhar, Yujin Kim, Scott Counts, Jennifer Neville, Siddharth Suri, Chirag Shah, Ryen W. White, Longqi Yang, Reid Andersen, Georg Buscher, Dhruv Joshi, Nagu Rangan, “TnT-LLM: Text Mining at Scale with Large Language Models”, in KDD 2024, https://dl.acm.org/doi/pdf/10.1145/3637528.3671647
  • Siru Ouyang, Jiaxin Huang, Pranav Pillai, Yunyi Zhang, Yu Zhang, Jiawei Han, “Ontology Enrichment for Effective Fine-grained Entity Typing”, in KDD’24 https://doi.org/10.48550/arXiv.2310.07795

Section 2

Section 3

Section 4

  • Zhaoheng Li, Supawit Chockchowwat, Ribhav Sahu, Areet Sheth, Yongjoo Park. “Kishu: Time-Traveling for Computational Notebooks.” arxiv’24. https://arxiv.org/abs/2406.13856
  • Devin Petersohn, Stephen Macke, Doris Xin, William Ma, Doris Lee, Xiangxi Mo, Joseph E. Gonzalez, Joseph M. Hellerstein, Anthony D. Joseph, Aditya Parameswaran. “Towards scalable dataframe systems.” PVLDB’20 https://arxiv.org/pdf/2001.00888

Section 5  

  • Jiang, Pengcheng, Cao Xiao, Zifeng Wang, Parminder Bhatia, Jimeng Sun, and Jiawei Han. 2024. “TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale.” arXiv [Cs.CL]. arXiv. http://arxiv.org/abs/2403.10351.
  • Wang, Hanyin, Chufan Gao, Christopher Dantona, Bryan Hull, and Jimeng Sun. 2024. “DRG-LLaMA : Tuning LLaMA Model to Predict Diagnosis-Related Group for Hospitalized Patients.” NPJ Digital Medicine 7 (1): 16. https://arxiv.org/abs/2309.12625

Section 6

Section 7

  • Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024. The Power of Noise: Redefining Retrieval for RAG Systems. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 719–729. https://doi.org/10.1145/3626772.3657834
  • Alireza Salemi, Surya Kallumadi, and Hamed Zamani. 2024. Optimization Methods for Personalizing Large Language Models through Retrieval Augmentation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 752–762. https://doi.org/10.1145/3626772.3657783

Section 8

  •  Lyu, Hanjia, et al. “Llm-rec: Personalized recommendation via prompting large language models.” arXiv preprint arXiv:2307.15780 (2023). https://arxiv.org/abs/2307.15780
  • Ye, Ruosong, et al. “Language is all a graph needs.” Findings of the Association for Computational Linguistics: EACL 2024. 2024. https://arxiv.org/abs/2308.07134

Section 9

Section 10

  • Nishant Balepur, Jie Huang, Kevin Chen-Chuan Chang: Expository Text Generation: Imitate, Retrieve, Paraphrase. EMNLP 2023: 11896-11919. https://arxiv.org/abs/2305.03276
  • DEER: Descriptive Knowledge Graph for Explaining Entity Relationships. Jie Huang, Kerui Zhu, Kevin Chen-Chuan Chang, Jinjun Xiong, Wen-mei Hwu. In The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022. PDF file: https://arxiv.org/abs/2205.10479
  • Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9248–9274, Singapore. Association for Computational Linguistics. https://aclanthology.org/2023.findings-emnlp.620/

Section 11

Fall 2024 Reading List

The written exam of the DAIS Qual Exam in Fall 2024 will be held on Friday,  Oct. 4, 2024, at 3pm-7pm in room 0220 Siebel Center (tentative), which is in the basement of Siebel Center (floor map: https://facilityaccessmaps.fs.illinois.edu/archibus/schema/ab-products/essential/workplace/?blId=0563&flId=00). 

This reading list consists of  multiple topic sections, each containing 2-3 papers.  The questions in the written exam will be based on the papers listed here, with 1-2 questions related to each section. If a section has two papers, you can usually expect to see one question related to the section in the qual exam, while if a section has three papers, you can usually expect to see two questions related to the section. You only need to answer four of those questions in the exam, so there is no need for you to read every paper. Instead, it would make sense for you to browse through the list and identify up to 4 sections that have papers that you are most familiar with or most comfortable with reading, and then focus on reading/digesting those papers. In general, you will likely find some sections to be closer to your interests or background than others, and you can focus more on reading the papers in those a few sections that seem to be closest to your research interests.    

Section 1

  • Mengting Wan, Tara Safavi, Sujay Kumar Jauhar, Yujin Kim, Scott Counts, Jennifer Neville, Siddharth Suri, Chirag Shah, Ryen W. White, Longqi Yang, Reid Andersen, Georg Buscher, Dhruv Joshi, Nagu Rangan, “TnT-LLM: Text Mining at Scale with Large Language Models”, in KDD 2024, https://dl.acm.org/doi/pdf/10.1145/3637528.3671647
  • Siru Ouyang, Jiaxin Huang, Pranav Pillai, Yunyi Zhang, Yu Zhang, Jiawei Han, “Ontology Enrichment for Effective Fine-grained Entity Typing”, in KDD’24 https://doi.org/10.48550/arXiv.2310.07795

Section 2

Section 3

Section 4

  • Zhaoheng Li, Supawit Chockchowwat, Ribhav Sahu, Areet Sheth, Yongjoo Park. “Kishu: Time-Traveling for Computational Notebooks.” arxiv’24. https://arxiv.org/abs/2406.13856
  • Devin Petersohn, Stephen Macke, Doris Xin, William Ma, Doris Lee, Xiangxi Mo, Joseph E. Gonzalez, Joseph M. Hellerstein, Anthony D. Joseph, Aditya Parameswaran. “Towards scalable dataframe systems.” PVLDB’20 https://arxiv.org/pdf/2001.00888

Section 5  

  • Jiang, Pengcheng, Cao Xiao, Zifeng Wang, Parminder Bhatia, Jimeng Sun, and Jiawei Han. 2024. “TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale.” arXiv [Cs.CL]. arXiv. http://arxiv.org/abs/2403.10351.
  • Wang, Hanyin, Chufan Gao, Christopher Dantona, Bryan Hull, and Jimeng Sun. 2024. “DRG-LLaMA : Tuning LLaMA Model to Predict Diagnosis-Related Group for Hospitalized Patients.” NPJ Digital Medicine 7 (1): 16.

Section 6

Section 7

  • Florin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice, Cesare Campagnano, Yoelle Maarek, Nicola Tonellotto, and Fabrizio Silvestri. 2024. The Power of Noise: Redefining Retrieval for RAG Systems. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 719–729. https://doi.org/10.1145/3626772.3657834
  • Alireza Salemi, Surya Kallumadi, and Hamed Zamani. 2024. Optimization Methods for Personalizing Large Language Models through Retrieval Augmentation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 752–762. https://doi.org/10.1145/3626772.3657783

Section 8

  •  Lyu, Hanjia, et al. “Llm-rec: Personalized recommendation via prompting large language models.” arXiv preprint arXiv:2307.15780 (2023). https://arxiv.org/abs/2307.15780
  • Ye, Ruosong, et al. “Language is all a graph needs.” Findings of the Association for Computational Linguistics: EACL 2024. 2024. https://arxiv.org/abs/2308.07134

Section 9

Section 10

  • Nishant Balepur, Jie Huang, Kevin Chen-Chuan Chang: Expository Text Generation: Imitate, Retrieve, Paraphrase. EMNLP 2023: 11896-11919. https://arxiv.org/abs/2305.03276
  • DEER: Descriptive Knowledge Graph for Explaining Entity Relationships. Jie Huang, Kerui Zhu, Kevin Chen-Chuan Chang, Jinjun Xiong, Wen-mei Hwu. In The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022. PDF file: https://arxiv.org/abs/2205.10479
  • Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9248–9274, Singapore. Association for Computational Linguistics. https://aclanthology.org/2023.findings-emnlp.620/

Section 11

Towards Theoretical Understanding of Overparametrization in Deep Learning

Guest Speaker: Jason Lee, Assistant Professor, University of Southern California

Date/Time: Thursday, October 4, 2018, 2:00 pm

Location:  3403 Siebel

Sponsored by the Department of Computer Science and DAIS

 

Abstract:  We provide new theoretical insights on why over-parametrization is effective in learning neural networks. For a k hidden node shallow network with quadratic activation and n training data points, we show as long as k>=sqrt(2n), over-parametrization enables local search algorithms to find a \emph{globally} optimal solution for general smooth and convex loss functions. Further, despite that the number of parameters may exceed the sample size, we show with weight decay, the solution also generalizes well.

Next, we analyze the implicit regularization effects of various optimization algorithms. In particular we prove that for least squares with mirror descent, the algorithm converges to the closest solution in terms of the bregman divergence. For linearly separable classification problems, we prove that the steepest descent with respect to a norm solves SVM with respect to the same norm. For over-parametrized non-convex problems such as matrix sensing or neural net with quadratic activation, we prove that gradient descent converges to the minimum nuclear norm solution, which allows for both meaningful optimization and generalization guarantees.

 

Bio: Jason Lee is an assistant professor in Data Sciences and Operations at the University of Southern California. Prior to that, he was a postdoctoral researcher at UC Berkeley working with Michael Jordan. Jason received his PhD at Stanford University advised by Trevor Hastie and Jonathan Taylor. His research interests are in statistics, machine learning, and optimization. Lately, he has worked on high dimensional statistical inference, analysis of non-convex optimization algorithms, and theory for deep learning.

This is joint work with Suriya Gunasekar, Mor Shpigel, Daniel Soudry, Nati Srebro, and Simon Du.

Xiang Ren, PhD Candidate, “Minimal-Effort StructMine: Turning Massive Corpora into Structures”

Time: 11am-12pm

Location:  3401 Siebel Center

Abstract:

While “Big Data” technologies are gaining great successes in unlocking knowledge from structured data, real-world data are largely unstructured and in the form of natural-language text. One of the grand challenges is to turn such massive text data into machine-actionable structures. Yet, most existing systems have heavy reliance on human efforts when dealing with text corpora of various kinds, slowing down the development of downstream applications.

In this talk, I will introduce a data-driven framework, minimal-effort StructMine, that extracts factual structures from massive text corpora with minimal human involvement. In particular, I will discuss how to apply Minimal-Effort StructMine to solve three subtasks: from identifying typed entities in text, to refining entity types into more fine-grained levels, to understanding the typed relationships between entities. Together, these three solutions form a clear roadmap for turning a massive corpus into a structured network to represent factual knowledge. Finally, I will share some directions towards mining corpus-specific structured networks for knowledge discovery.

Bio:

Xiang Ren is a Computer Science PhD candidate at University of Illinois at Urbana-Champaign, working with Jiawei Han and the Data and Information System Lab. Xiang’s research develops data-driven methods for turning unstructured text data into machine-actionable structures. More broadly, his research interests span data mining, machine learning, and natural language processing, with a focus on making sense of massive text corpora. His research has been recognized with a Google PhD Fellowship, Yahoo!-DAIS Research Excellence Award, C. W. Gear Outstanding Graduate Student Award, and has been transferred to US Army Research Lab and Microsoft Bing.

Data and Intelligent Systems
Email: yongjoo@g.illinois.edu