I welcome motivated Master's students interested in research-oriented projects in Big Data, distributed data systems, cloud computing, optimization, and AI-powered data systems. Projects are designed to combine technical implementation with careful experimentation, evaluation, and scholarly communication.
A strong project should go beyond implementing an existing tutorial. The goal is to identify a meaningful problem, build a sound system or method, evaluate it systematically, and communicate the results clearly.
A clearly defined problem with measurable objectives, meaningful comparisons, and room for investigation.
A working implementation using appropriate data, algorithms, platforms, and scalable computing technologies.
Experiments, metrics, analysis, documentation, and a final report or paper-quality presentation of the findings.
These areas reflect my established expertise and current directions. Specific project topics can be refined based on a student's preparation, interests, and available datasets or computing resources.
Projects involving large-scale data processing, distributed analytics, scientific workflows, and scalable data management.
Apache SparkSpark SQLDistributed ProcessingBig Data AnalyticsProjects exploring how modern AI can be integrated with large-scale data systems, including retrieval, semantic search, and intelligent data analytics.
RAGLLMsVector SearchAI + Big DataData placement, resource allocation, workflow scheduling, and optimization for data-intensive cloud applications.
Cloud ComputingSchedulingOptimizationResource AllocationData-driven and AI-assisted approaches to workload scheduling, energy efficiency, thermal management, and computing-resource optimization.
Data CentersEnergy EfficiencyWorkload SchedulingAnalyticsThe examples below are starting points rather than fixed assignments. A Master's project should be narrowed into a specific research question before implementation begins.
Investigate architectures for preprocessing, indexing, retrieving, and querying large document collections using distributed data processing and vector retrieval.
Evaluate indexing strategies, retrieval quality, latency, scalability, and metadata filtering for large vector collections.
Explore natural-language interfaces or AI-assisted methods for querying, interpreting, and analyzing large structured or unstructured datasets.
Study the use of LLMs or related AI techniques for data preparation, metadata generation, data-quality analysis, or pipeline assistance.
Develop and evaluate scheduling, placement, or resource-allocation approaches for data-intensive workflows in cloud environments.
Use workload and system data to investigate thermal-aware, energy-efficient, or resource-aware scheduling and prediction.
Students do not need to know every technology before starting. The exact stack depends on the project, but a solid programming and data-systems foundation is important.
If you are interested in working with me, the most useful first message is concise and specific. You do not need to arrive with a complete research proposal.
Identify one or two research areas or example directions above that genuinely interest you.
Tell me your degree program, relevant courses or experience, technical strengths, and the area you would like to explore.
We can refine the idea based on research value, prerequisites, available data, computing resources, and your timeline.
A project normally progresses through literature review, problem definition, system design, implementation, experiments, analysis, and final writing.
If your interests align with Big Data, distributed data systems, cloud computing, optimization, or AI-powered data systems, I welcome a short introduction describing your background and the direction you would like to explore.