Wanqi Yang
I focus on AI models, systems, hardware, and the productization trends that connect them. My long-term vision is to help build a more individually sovereign digital future by making machine learning systems more efficient, stable, and accessible.
I try to work toward this goal in three ways: as a researcher, by pushing forward questions in efficient machine learning systems; as a developer, by building tools and systems in this space; and as a writer, by using my personal blog to share knowledge, experience, and practical intuition about efficient ML systems and the latest AI tools.
As a researcher
My research goal is to advance foundation-model systems from multiple perspectives: improving model capabilities, developing deeper understanding of model behavior, and making ML systems more efficient, reliable, and accessible. I like to take a whole-stack view, from representations and learning algorithms down to inference systems and low-level hardware backends.
Currently, I work with Dr. Shiwei Liu and Dr. Yuexiao Ma on empirical studies for improving the efficiency of foundation models. One result from this line of work is AlphaQ, a calibration-free bit-allocation method for Mixture-of-Experts quantization. More recently, we have been focusing on inference optimization for multimodal foundation models.
Before this, I completed my master's thesis in the Hardware for Artificial Intelligence Lab at TU Darmstadt, where I studied classification-difficulty-aware neural network quantization. I also worked as a research assistant in the Artificial Intelligence & Machine Learning Lab at TU Darmstadt, building a cloud-based experimental environment for studying learning-agent behavior through multi-agent interactions. My undergraduate thesis was completed in the robotics lab at Beijing Jiaotong University, where I carried out the full design process for a multi-form robot and received an Outstanding Undergraduate Thesis award.
As a developer
In my engineering work, I have built and optimized industrial-scale inference systems where model quality, latency, memory, operators, and hardware constraints all have to be handled together.
At Qualcomm, I work as a Machine Learning Engineer on multimodal model inference for mobile and PC platforms. My work involves adapting foundation models for Snapdragon platforms across model structure, quantization, operators, speculative decoding, and deployment pipelines.
Earlier at DJI, I worked on heterogeneous inference for autonomous-driving perception models, including preprocessing, custom operators, kernels, compiler passes, and hardware-aware execution across Qualcomm CPU, GPU, and DSP backends. This background keeps pulling me toward a whole-stack view: representations and learning algorithms on one side, deployment pipelines and low-level backends on the other.
I also like building small tools that make a workflow more inspectable or reduce the friction between an idea and a runnable system. Recently, I have been especially interested in building local AI products in the macOS ecosystem. A few of these experiments are collected on my Project page.
As a writer
I write in Chinese under the name Vinci叽里呱啦. I treat it as a habit of learning and self-expression. I mainly record how I understand complex concepts that interest me, along with my own practical experience. I publish on Zhihu, WeChat Official Account, and X; you are warmly welcome to follow along.
How I got here
My path into this area has been driven by curiosity. As an undergraduate at Beijing Jiaotong University, I used my laptop for computation and simulation work related to robotics and racing-car aerodynamics. That experience made the value of computing power very concrete to me: faster and more accessible computation changes how quickly we can iterate, test, and explore in science and engineering. It also led me to pursue a master's degree in Computational Engineering at TU Darmstadt.
Outside work
I spend a lot of time reading, playing tennis, hiking, and taking photographs. Reading keeps me close to longer histories and slower arguments; tennis and hiking give me a more physical rhythm away from the screen; photography trains a different kind of attention, one that is less about optimizing a system and more about noticing what is already there. I am also a tarot reader; as a game of symbols, tarot helps me keep a more intuitive sense of how meanings can connect.