We are looking for AI Infrastructure Engineers to build the data infrastructure, AI development platform, MLOps/LLMOps toolchains, model serving systems, and agent infrastructure;
The Engineers in this role will work closely with applied AI engineers, scientists and domain experts to understand concrete research bottlenecks and translate them into reliable, reusable and scalable AI engineering solutions.
This role spans three complementary tracks. You are not expected to cover all of them. Each engineer will take primary ownership of ONE to TWO tracks based on their experience, strengths, and team needs.
我们正在招聘 AI 基础设施工程师,负责建设现代 AI 工程及研发团队所需的数据基础设施、AI 开发平台、MLOps/LLMOps 工具链、模型基础设施和智能体基础设施。
该岗位的工程师将与AI应用工程师、科学家、领域专家紧密合作,理解具体科研瓶颈,并将其转化为可靠、可复用、可扩展的AI工程解决方案。
该岗位涵盖三个相互协同的方向。我们不期待任何一位候选人同时覆盖所有方向。*每位工程师将根据自身经验、优势和团队需求,主要负责其中一到两个方向。
Responsibilities | 职责:
Track 1: Data Engineering
方向一:数据工程
- Design and build scalable data pipelines for AI model development, training, evaluation, and production applications.
设计并建设可扩展的数据 Pipeline,支持 AI 模型开发、训练、评估和生产应用。
- Build ingestion, transformation, validation, versioning, and lineage capabilities for structured and unstructured data.
建设结构化与非结构化数据的采集、转换、校验、版本管理和数据血缘能力。
- Develop reliable workflows for processing text, images, documents, scientific data, time-series data, and other domain-specific datasets.
开发可靠的数据处理流程,支持文本、图像、文档、科学数据、时序数据及其他领域数据。
- Build data infrastructure for LLM and RAG applications, including vector stores, knowledge bases, chunking and embedding pipelines, and retrieval quality evaluation.
建设面向大模型与 RAG 应用的数据基础设施,包括向量存储、知识库、切分与向量化 Pipeline,以及检索质量评估。
- Improve data quality, observability, reproducibility, and accessibility across AI development workflows.
提升 AI 开发流程中的数据质量、可观测性、可复现性和可访问性。
- Work with algorithm and domain teams to translate model development requirements into reusable data infrastructure.
与算法团队和业务领域团队合作,将模型开发需求转化为可复用的数据基础设施。
Track 2: AI Platform, MLOps/LLMOps and Model Infrastructure
方向二:AI 平台、MLOps/LLMOps 与模型基础设施
- Design and build AI development platforms that support the end-to-end model lifecycle, including data preparation, experimentation, training, evaluation, deployment, inference, monitoring, and iteration.
设计并建设支持模型全生命周期的 AI 开发平台,包括数据准备、实验、训练、评估、部署、推理、监控和持续迭代。
- Build standardized MLOps and LLMOps pipelines, tools, and workflows for model development, lifecycle management, and production delivery.
建设标准化的 MLOps 和 LLMOps Pipeline、工具及工作流,支持模型开发、全生命周期管理和生产交付。
- Provide reusable platform capabilities for experiment tracking, dataset and model versioning, model evaluation, model registry, release management, and reproducibility.
提供可复用的平台能力,支持实验追踪、数据集与模型版本管理、模型评估、模型注册、发布管理和结果复现。
- Build scalable training, fine-tuning, evaluation, and inference workflows across local, cloud, Kubernetes, GPU cluster, and hybrid environments.
建设适用于本地、云端、Kubernetes、GPU 集群及混合环境的可扩展训练、微调、评估和推理工作流。
- Develop standardized model deployment and serving capabilities, including model packaging, rollout, request routing, load balancing, autoscaling, monitoring, and failure recovery.
开发标准化的模型部署与服务能力,包括模型封装、发布、请求路由、负载均衡、弹性扩缩、监控和故障恢复。
- Integrate and operate mainstream model training and serving frameworks such as PyTorch, DeepSpeed, FSDP, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or equivalent technologies.
集成并运行主流模型训练和服务框架,例如 PyTorch、DeepSpeed、FSDP、vLLM、SGLang、TensorRT-LLM、Triton Inference Server 或同类技术。
- Optimize model training and inference systems for performance, scalability, GPU utilization, memory efficiency, reliability, and cost.
围绕性能、可扩展性、GPU 利用率、显存效率、可靠性和成本,优化模型训练与推理系统。
- Improve AI developer experience by providing standardized development environments, SDKs, APIs, templates, CI/CD pipelines, and self-service platform capabilities.
通过标准化开发环境、SDK、API、模板、CI/CD Pipeline 和自助式平台能力,提升 AI 开发者体验。
- Build observability, resource scheduling, capacity management, cost monitoring, security, and governance capabilities for AI workloads.
建设 AI 工作负载的可观测性、资源调度、容量管理、成本监控、安全及治理能力。
- Collaborate with research, data, and application teams to translate model development requirements into reusable platform and infrastructure capabilities.
与研发、数据及应用团队协作,将模型开发需求转化为可复用的平台与基础设施能力。
Track 3: Agent Infrastructure
方向三:智能体基础设施
- Design and build infrastructure for developing, deploying, operating, and evaluating AI agents.
设计并建设支持 AI 智能体开发、部署、运行和评估的基础设施。
- Build reusable agent runtime capabilities, including model access, tool invocation, workflow orchestration, state management, memory, and execution control.
建设可复用的 Agent Runtime 能力,包括模型接入、工具调用、工作流编排、状态管理、记忆和执行控制。
- Develop infrastructure for agent identity, permissions, sandboxing, secrets management, observability, tracing, and auditability.
建设智能体身份、权限、沙箱、密钥管理、可观测性、链路追踪和审计能力。
- Build agent evaluation and testing systems covering task completion, tool-use correctness, reliability, safety, latency, and cost.
建设智能体评估与测试系统,覆盖任务完成度、工具调用正确性、可靠性、安全性、延迟和成本。
- Provide reusable SDKs, APIs, templates, skills, connectors, and deployment workflows for agent development teams.
为智能体开发团队提供可复用的 SDK、API、模板、Skills、连接器和部署工作流。
- Support long-running, event-driven, scheduled, asynchronous and human-in-the-loop agent workflows.
支持长时间运行、事件驱动、定时执行、异步执行和以及人在回路中的智能体工作流。
- Integrate agents with enterprise systems, data platforms, model services, internal tools, and external APIs, including standardized protocols such as MCP.
将智能体与企业系统、数据平台、模型服务、内部工具及外部 API 进行集成,包括 MCP 等标准化协议。
Qualifications | 任职要求:
Bachelor’s degree or above, majoring in Computer Science, Software Engineering, Cloud Computing, Big Data, System Engineering, Math,Artificial Intelligence or related fields.
本科及以上学历,计算机科学、软件工程、云计算、大数据、系统工程、数学、人工智能等相关专业
Strong programming and software engineering skills in Python, Java, Go, Scala, C++, Rust, or another relevant programming language.
具备扎实的编程和软件工程能力,熟悉 Python、Java、Go、Scala、C++、Rust 或其他相关编程语言。
Familiarity with relevant technologies in at least one professional track, such as data engineering, workflow orchestration systems, cloud platforms, Docker, Kubernetes, GPU clusters, distributed training frameworks, model-serving systems, or agent frameworks.
熟悉至少一个专业方向的相关技术,例如数据工程、工作流编排系统、云平台、Docker、Kubernetes、GPU 集群、分布式训练框架、模型服务系统或 Agent 框架。
Ability to translate research, algorithm, data, or business requirements into reusable engineering platforms, infrastructure, tools, and workflows.
能够将研究、算法、数据或业务需求转化为可复用的工程平台、基础设施、工具和工作流。
Strong system design, problem-solving, debugging, and performance-analysis capabilities.
具备较强的系统设计、问题解决、故障排查和性能分析能力。
Strong communication and cross-functional collaboration skills, with the ability to work effectively with research, algorithm, data, product, and application teams.
具备良好的沟通和跨团队协作能力,能够与研究、算法、数据、产品及应用团队高效合作。
Preferred Qualifications | 优先条件:
3+years of experience designing, building, deploying, or operating reliable production systems.
具备3年以上设计、建设、部署或运行高可用生产系统的经验。
Experience with AI agent frameworks, workflow engines, tool-use systems, memory systems, sandbox environments, or agent evaluation methods.
具备 AI Agent 框架、工作流引擎、工具调用系统、记忆系统、沙箱环境或智能体评估经验。
Experience supporting AI workloads in cloud, hybrid-cloud, on-premises, or multi-cluster environments.
具备在云端、混合云、本地或多集群环境中支持 AI 工作负载的经验。
Experience optimizing data systems for performance, scalability and cost
具备数据系统的性能、可扩展性与成本优化经验。
Experience building AI infrastructure for scientific computing, materials science, energy, manufacturing, or other domain-specific applications.
具备为科学计算、材料科学、能源、制造或其他领域应用建设 AI 基础设施的经验。
Experience adapting large models to specialized domains(scientific computing materials, energy, industrial applications)
具备将大模型适配到专业领域(科学计算、材料、能源、工业应用)的经验。
Open-source contributions, track record of academic publications, granted patents or participation in industry standards.
具备开源贡献、学术发表记录,已授权专利或行业标准参与经历。