Bohan Hou

I am a final-year Ph.D. student in Computer Science at Carnegie Mellon University, where I am very fortunate to be advised by Prof. Tianqi Chen. I am a member of Catalyst.

My research interests lie in machine learning systems and deep learning compilers.

Prior to coming to CMU, I received my Bachelor's degree in ACM Honors Class of Shanghai Jiao Tong University, where I worked under the supervision of Prof. Weinan Zhang and Prof. Yong Yu.

GitHub / LinkedIn / Google Scholar

Photo at Badwater, Death Valley
Photo taken at Badwater Basin, Death Valley National Park.
Projects

TIRx: An Open Compiler Stack for Evolving Frontier ML Kernels [post]
TIRx is an open-source, hardware-native DSL and compiler for ML kernels, built on Apache TVM. It compiles to GPUs and specialized AI accelerators today and is designed to grow with future hardware generations. The same design serves expert-written kernels, agent-generated kernels, and megakernel systems.

TIRx Harness: An Open Compiler Harness for Agentic GPU Programming [kernels] [harness] [post]
TIRx Harness is an open compiler harness for agentic GPU programming. 2.94× KDA fwd, 6.84× KDA bwd, 2.59× MSA prefill, 3.99× MSA decode, 1.33× KDA decode, 1.71× MLA, 1.68× VSA — just by pressing Enter.

MLC-LLM: Machine Learning Compilation for Large Language Models [repo]
MLC-LLM is a universal solution that allows any language models to be deployed natively on a diverse set of hardware backends and native applications, plus a productive framework for everyone to further optimize model performance for their own use cases.

Research (*indicates equal contribution)

TIRx: A Unified Tile Primitive ML Compiler Across Heterogeneous Backends [pdf]
Bohan Hou, Hongyi Jin, Guanjie Wang, Shushi Hong, Jinqi Chen, Yaxing Cai, Lijie Yang, Zihao Ye, Yaoyao Ding, Ruihang Lai, Todd Mowry, Tianqi Chen
The 32nd ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS '27)

Gecko: An Efficient Neural Architecture Inherently Processing Sequences with Arbitrary Lengths [pdf] [openreview]
Xuezhe Ma, Shicheng Wen, Linghao Jin, Bilge Acun, Ruihang Lai, Bohan Hou, Will Lin, Hao Zhang, Songlin Yang, Ryan Lee, Mengxi Wu, Jonathan May, Luke Zettlemoyer, Carole-Jean Wu
Conference on Language Modeling (COLM '26)

CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution [pdf]
Zihao Ye, Yingyi Huang, Hongyi Jin, Bohan Hou, Junru Shao, Zhongming Yu, Jinqi Chen, Meghan Cowan, Shiyi Cao, Shanli Xing, Hanfeng Chen, Vinod Grover, Tianqi Chen, Luis Ceze
Preprint, arXiv:2608.12629 (2026)

MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs [pdf]
Xinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji, Jinchen Jiang, Zepeng Zhao, Ziruo Xiao, Zihao Ye, Yingyi Huang, Ruihang Lai, Hongyi Jin, Bohan Hou, Mengdi Wu, Yixin Dong, Anthony Yip, Zihao Ye, Songting Wang, Wenqin Yang, Xupeng Miao, Tianqi Chen, Zhihao Jia
The 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI '26)

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel [pdf]
Hongyi Jin*, Bohan Hou*, Guanjie Wang*, Ruihang Lai*, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C. Mowry, Zhihao Jia, Tianqi Chen
Conference on Machine Learning and Systems (MLSys '26)

Tilus: A Virtual Machine for Arbitrary Low-Precision GPGPU Computation in LLM Serving [pdf]
Yaoyao Ding, Bohan Hou, Xiao Zhang, Allan Lin, Tianqi Chen, Cody Yu Hao, Yida Wang, Gennady Pekhimenko
The 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS '26)

Relax: Composable Abstractions for End-to-End Dynamic Machine Learning [pdf]
Ruihang Lai*, Junru Shao*, Siyuan Feng*, Steven S Lyubomirsky*, Bohan Hou, Wuwei Lin, Zihao Ye, Hongyi Jin, Yuchen Jin, Jiawei Liu, Lesheng Jin, Yaxing Cai, Ziheng Jiang, Yong Wu, Sunghyun Park, Prakalp Srivastava, Jared G Roesch, Todd C Mowry, Tianqi Chen
The 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS '25)

WebLLM: A High-Performance In-Browser LLM Inference Engine [pdf]
Charlie F. Ruan, Yucheng Qin, Akaash R. Parthasarathy, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, Tianqi Chen
arXiv preprint arXiv:2412.15803 (2024)

Optimal Kernel Orchestration for Tensor Programs with Korch [pdf]
Muyan Hu*, Ashwin Venkatram*, Shreyashri Biswas*, Balamurugan Marimuthu*, Bohan Hou, Gabriele Oliaro, Haojie Wang, Liyan Zheng, Xupeng Miao, Jidong Zhai, Zhihao Jia
The 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS '24)

Accelerating Self-Attentions for LLM Serving with FlashInfer [post]
Zihao Ye, Lequn Chen, Ruihang Lai, Yilong Zhao, Size Zheng, Junru Shao, Bohan Hou, Hongyi Jin, Yifei Zuo, Liangsheng Yin, Tianqi Chen, Luis Ceze
FlashInfer blog post (2024)

TensorIR: An Abstraction for Automatic Tensorized Program Optimization [pdf]
Siyuan Feng*, Bohan Hou*, Hongyi Jin, Wuwei Lin, Junru Shao, Ruihang Lai, Zihao Ye, Lianmin Zheng, Cody Hao Yu, Yong Yu, and Tianqi Chen
The 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS '23)

Tensor Program Optimization with Probabilistic Programs [pdf]
Junru Shao, Xiyou Zhou, Siyuan Feng, Bohan Hou, Ruihang Lai, Hongyi Jin, Wuwei Lin, Masahiro Masuda, Cody Hao Yu, Tianqi Chen
Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS), 2022