Wei Yang

Full-Stack Software Engineer at PiSrc

PiSrc 全栈软件工程师

Email:邮箱: hey.weiyang@gmail.com

Phone:电话: +1 551 260 0541

Web:网站: inscribedeeper.github.io

Professional Summary

Full-stack software engineer and data science practitioner with experience in production generative AI, enterprise search, natural language processing, recommendation systems, and data integration. Work spans secure RAG applications, multilingual retrieval pipelines, API-based enterprise data synchronization, cloud infrastructure, and applied NLP research. Experienced in translating machine learning methods into reliable systems for industrial and business platforms.

Areas of Expertise

Generative AI & Information Retrieval: Retrieval-augmented generation, conversational AI, agentic workflows, semantic search, hybrid retrieval, reranking, and AI safety controls

Enterprise Platforms & Data: Multilingual indexing, knowledge management, API integration, Elasticsearch, Solr, Weaviate, Redis, MuleSoft, Adobe Experience Manager, Azure, and AWS

Software Engineering: Python, Java, JavaScript, SQL, REST APIs, distributed caching, load balancing, Docker, Nginx, monitoring, and performance optimization

Machine Learning & Analytics: PyTorch, Hugging Face Transformers, scikit-learn, Pandas, NLP, recommendation systems, text classification, and statistical modeling

Experience

PiSrc

Full-Stack Software Engineer

Feb 2022 – Present · New York, NY

Production AI, search, CMS, and personalization for industrial digital platforms.

  • Built production RAG conversational AI and hybrid-search experiences with Azure OpenAI, Redis, and Weaviate
  • Designed recommendation and offline ML pipelines for personalization, funnel optimization, and targeting
  • Integrated partner/location CRM updates via MuleSoft into Solr and Elasticsearch with caching
  • Developed licensed AEM CMS software for content authoring and publishing, integrating with search, PIM, marketing, and ecommerce platforms
  • Translated business and UX requirements into modular, decoupled components; maintained data stores, documentation, and QA for reliable operations
  • Hardened delivery with secure AI workflows, multilingual retrieval, and maintainable cross-system data flows

Stevens Institute of Technology

Research Assistant

Jun 2020 – Dec 2021 · Hoboken, NJ

Deep learning and NLP engineering for financial text analysis.

  • Built and maintained data-processing and training workflows for NLP experiments on financial text
  • Modularized pipelines and PyTorch/ML helpers into reusable utils for shared research use
  • Set up and operated remote Ubuntu GPU environments for deep learning training and evaluation
  • Supported language-feature and text-classification studies with domain-adapted BERT models

Industry Projects

AI content transformation — PDFs become semantic HTML and structured AEM components for enterprise publishing.

  • Delivered PDF-to-semantic-HTML workflows that produce authored Adobe Experience Manager components
  • Implemented an agentic LLM-vision pipeline to parse layout and generate reusable AEM components
  • Added I18N translation support governed by brand voice and editorial guidelines
  • Designed the tools to run on existing customer infrastructure without external processing servers

Prism - PoC

2026 · Active · Core Contributor

Agentic AI conversation platform — knowledge-grounded engagement with playbooks and tools for marketing-funnel optimization.

  • Contributed to proof-of-concept development and architecture discussions for a platform that answers visitor questions from unified knowledge sources with citations
  • Provided architecture recommendations for knowledge ingestion, retrieval, agentic playbooks, internal tools, conversation flows, and enterprise deployment
  • Researched and optimized retrieval quality, response grounding, agentic routing, playbook and tool behavior, and marketing-funnel performance

AI Overview

2025 · Active · Lead

Real-time AI search overview — hybrid retrieval across heterogeneous enterprise sources.

  • Built real-time search overview and autosuggest with live-query caching for low-latency responses
  • Engineered multilingual indexing across 10+ sources with hybrid search, reranking, routing, and iterative retrieval
  • Added reporting pipelines and feedback loops to improve overview relevance and quality over time
  • Separated overview generation from chat so each surface could optimize latency, ranking, and UX independently

ABM Personalization

2025 · Maintenance

Account-based marketing — website experiences tailored to target accounts and buying contexts.

  • Extended recommendation workflows from product-level targeting to account-aware journeys
  • Aligned page content and offers with account segments and prior interaction signals
  • Connected personalization outputs to account segments used across marketing surfaces

AI Navigator

2024 · Active · Lead

Agentic conversational AI — enterprise knowledge retrieval with measurable production usage.

  • Built an agentic chatbot with Azure OpenAI, Redis, and Weaviate, integrating internal company tools and APIs through function calling
  • Designed multi-turn Redis memory and agentic function-calling workflows for tools and APIs
  • Implemented security guardrails including input sanitization, content filtering, and PII masking
  • March 2026 usage: approximately 110K unique sessions across the tracked search, retrieval-assistant, and fault-code applications
  • Delivered the website chat entry point for AI-powered support and quick access to trusted information

Personalization

2024 · Maintenance · Lead

Live demo在线演示: 123

Recommendations and offline ML — behavioral data turned into personalized product experiences.

  • Built user-to-item and item-to-item recommendation pipelines with rolling cache and Akamai CDN
  • Designed offline ML pipelines for interaction modeling, personalization, and user classification
  • Improved sales-funnel conversion and marketing targeting through customized product experiences
  • Delivered discovery surfaces such as “Based on your views,” preferred products, and related content

Fault Code Navigator

2023 · Maintenance · Lead

AI-assisted technical support — fault-code lookup across industrial product documentation.

  • Built AI-assisted lookup of condition, event, and fault codes across industrial product families
  • Connected product-family selection and code queries to Technical Documentation Center retrieval
  • Delivered production troubleshooting access for drives, controllers, and motion systems

Content Score

2023 · Maintenance · Lead

Content-quality scoring — ML signals that accelerate marketing-funnel decisions.

  • Built a backend analytics platform that scores content quality for funnel optimization
  • Developed ML workflows to evaluate content effectiveness and prioritize conversion-focused changes
  • Connected scoring outputs with marketing and content workflows for faster data-driven decisions

Partner Locator

2022 · Active

CRM-to-search data integration — partner and location records made reliably discoverable.

  • Designed 7+ MuleSoft API workflows for incremental CRM delta updates of partner accounts and locations
  • Integrated multi-source records into Lucidworks Solr and Elasticsearch search with cache optimization
  • Improved reliability of partner and location data for searchable enterprise applications
  • Continues in active maintenance and enhancement for production partner-locator experiences

Assessment questionnaires in AEM — first-party preference capture for cybersecurity readiness.

  • Tailored AEM questionnaire components with category support, reporting, and dashboard presentation
  • Supported a Cybersecurity Preparedness Assessment that produces a NIST-based score and customized report
  • Improved structured authoring and multi-channel publishing for assessment-driven content workflows

Site Feedback

2022 · Maintenance

On-site feedback loop — Hotjar signals wired into AEM-managed pages.

  • Integrated Hotjar with Adobe Experience Manager to capture on-site feedback and behavioral signals
  • Connected feedback collection to AEM pages so teams could identify UX friction and prioritize fixes
  • Established an early feedback loop that informed later personalization and content-scoring work

Education

Monroe University

Master of Business Administration (MBA), Part-time · In Progress

2025 - Present

Selected Coursework: Strategic Marketing & Data Mining, Software System Design, Computer Networks, Research & Statistics for Managerial Decision-Making, Organizational Behavior & Leadership in the 21st Century, Managing in the Global Environment

Stevens Institute of Technology

MSc in Data Science

September 2019 - December 2021

Relevant Coursework: Statistical Methods, Statistical Inference, Advanced Optimization Methods, Advanced Data Analytics & Machine Learning, Deep Learning, Natural Language Processing, Web Analytics, Database Management Systems, Web Programming, Data Structures & Algorithms

Guangzhou University

BSc in Mathematics and Applied Mathematics

September 2014 - May 2018

Relevant Coursework: Probability and Mathematical Statistics, Operational Research, Numerical Analysis, Advanced Algebra, Mathematical Analysis, Real Function Theory, Functional analysis, Ordinary Differential Equations, Partial Differential Equations

Awards & Honors

Provost's Scholarship

Stevens Institute of Technology

2019

Awarded the Provost’s Scholarship.

CUMCM

Guangzhou University

2015

National Second Prize in the National Mathematical Modeling Contest; Provincial First Prize.

Extracurricular Activities

  • UBS Quant Hackathon — 2020
  • UBS Pitch Competition — 2020
  • Stevens HealthTech Hackathon — 2019

Selected Course Projects

Analysis on Factors Influencing Bitcoin

Stevens Institute of Technology

2021 · Hoboken, NJ

NLP and deep learning for tweet sentiment and market indicators.

  • Cleaned and analyzed 20M+ related tweets with sentiment features at second-level granularity
  • Engineered MACD/RSI-style indicators and back-tested multiple deep learning models for portfolio return

Quantify the AI impacts on Jobs skills

Stevens Institute of Technology

2021 · Hoboken, NJ

Job-description NLP with clustering and topic modeling.

  • Scraped and structured Indeed job-description text; filtered noise with K-means skillset clusters
  • Analyzed skill trends with topic modeling and aspect-extraction neural networks

E-commerce Recommender System

Stevens Institute of Technology

2020 · Hoboken, NJ

Collaborative filtering on user, item, and interaction data.

  • Built a recommender engine with memory-based and model-based collaborative filtering on JD.com interaction data
  • Evaluated on 7,000 user-item samples and reached 77% Top-10 accuracy

MyPlace Web Development

Stevens Institute of Technology

2020 · Hoboken, NJ

Full-stack web app for furniture and rental information exchange.

  • Led a four-person team to build a Node.js/Express app with MongoDB schemas and RESTful CRUD APIs
  • Implemented authentication, search, comments, and dashboard features for listing exchange

Fintech Pitch Competition

Stevens Institute of Technology

2020 · Hoboken, NJ

City vibrancy modeling from multi-source public data.

  • Built a vibrancy index from Google POIs, Instagram, Zillow, and BLS data with Pandas feature engineering
  • Fine-tuned XGBoost and Keras models to predict U.S. city prosperity trends

专业摘要

全栈软件工程师与数据科学实践者,具备生产级生成式 AI、企业搜索、自然语言处理、推荐系统与数据集成经验。工作覆盖安全 RAG 应用、多语言检索流水线、基于 API 的企业数据同步、云基础设施与应用 NLP 研究。善于将机器学习方法落地为面向工业与商业平台的可靠系统。

专长领域

生成式 AI 与信息检索: 检索增强生成、对话式 AI、智能体工作流、语义搜索、混合检索、重排序与 AI 安全控制

企业平台与数据: 多语言索引、知识管理、API 集成、Elasticsearch、Solr、Weaviate、Redis、MuleSoft、Adobe Experience Manager、Azure 与 AWS

软件工程: Python、Java、JavaScript、SQL、REST APIs、分布式缓存、负载均衡、Docker、Nginx、监控与性能优化

机器学习与分析: PyTorch、Hugging Face Transformers、scikit-learn、Pandas、NLP、推荐系统、文本分类与统计建模

工作经历

PiSrc

全栈软件工程师

2022年2月 – 至今 · New York, NY

面向工业数字平台的生产级 AI、搜索、CMS 与个性化。

  • 使用 Azure OpenAI、Redis 与 Weaviate 构建生产级 RAG 对话式 AI 与混合搜索体验
  • 设计用于个性化、漏斗优化与定向的推荐与离线 ML 流水线
  • 通过 MuleSoft 将合作伙伴/地点 CRM 更新集成到 Solr 与 Elasticsearch,并配合缓存
  • 开发可授权的 AEM CMS 软件,用于内容创作与发布,并与搜索、PIM、营销与电商平台集成
  • 将业务与 UX 需求转化为模块化、解耦组件;维护数据存储、文档与 QA,保障可靠运维
  • 通过安全 AI 工作流、多语言检索与可维护的跨系统数据流强化交付

史蒂文斯理工学院

研究助理

2020年6月 – 2021年12月 · Hoboken, NJ

面向金融文本分析的深度学习与 NLP 工程。

  • 构建并维护面向金融文本 NLP 实验的数据处理与训练工作流
  • 将流水线与 PyTorch/ML 辅助工具模块化为可复用 utils,供研究共享使用
  • 搭建并运维远程 Ubuntu GPU 环境,用于深度学习训练与评估
  • 以领域自适应 BERT 模型支持语言特征与文本分类研究

行业项目

AI 内容转换——将 PDF 转为语义 HTML 与结构化 AEM 组件,用于企业发布。

  • 交付 PDF 到语义 HTML 工作流,生成可编辑的 Adobe Experience Manager 组件
  • 实现智能体 LLM-vision 流水线以解析版面并生成可复用 AEM 组件
  • 增加受品牌语气与编辑规范约束的 I18N 翻译支持
  • 设计工具在客户现有基础设施上运行,无需外部处理服务器

Prism - PoC

2026 · 进行中 · 核心贡献者

智能体 AI 对话平台——知识 grounding 互动,配合 playbook 与工具优化营销漏斗。

  • 参与概念验证开发与架构讨论,支持基于统一知识源并附引用的访客问答
  • 就知识摄入、检索、智能体 playbook、内部工具、对话流程与企业部署提供架构建议
  • 研究并优化检索质量、回答 grounding、智能体路由、playbook 与工具行为,以及营销漏斗表现

AI Overview

2025 · 进行中 · 负责人

实时 AI 搜索概览——跨异构企业数据源的混合检索。

  • 构建实时搜索概览与自动建议,含实时查询缓存以实现低延迟响应
  • 工程化跨 10+ 数据源的多语言索引,含混合搜索、重排序、路由与迭代检索
  • 增加报表流水线与反馈闭环,持续改进概览相关性与质量
  • 将概览生成与聊天分离,使各界面可独立优化延迟、排序与 UX

ABM Personalization

2025 · 维护中

基于账户的营销(ABM)——按目标账户与购买情境定制网站体验。

  • 将推荐工作流从产品级定向扩展到账户感知旅程
  • 使页面内容与优惠与账户细分及既往交互信号对齐
  • 将个性化输出连接到跨营销界面使用的账户细分

AI Navigator

2024 · 进行中 · 负责人

智能体对话式 AI——具备可衡量生产使用量的企业知识检索。

  • 使用 Azure OpenAI、Redis 与 Weaviate 构建智能体聊天机器人,通过函数调用集成内部公司工具与 API
  • 设计多轮 Redis 记忆与面向工具/API 的智能体函数调用工作流
  • 实现安全护栏,包括输入净化、内容过滤与 PII 掩码
  • 2026 年 3 月使用量:在所追踪的搜索、检索助手与故障代码应用中约有 110K 独立会话
  • 交付网站聊天入口,提供 AI 驱动支持与可信信息的快捷访问

Personalization

2024 · 维护中 · 负责人

Live demo在线演示: 123

推荐与离线 ML——将行为数据转化为个性化产品体验。

  • 构建用户到物品与物品到物品推荐流水线,含滚动缓存与 Akamai CDN
  • 设计用于交互建模、个性化与用户分类的离线 ML 流水线
  • 通过定制化产品体验提升销售漏斗转化与营销定向
  • 交付如“Based on your views”、偏好产品与相关内容等发现界面

Fault Code Navigator

2023 · 维护中 · 负责人

AI 辅助技术支持——跨工业产品文档的故障代码查询。

  • 构建跨工业产品系列的条件、事件与故障代码 AI 辅助查询
  • 将产品系列选择与代码查询连接到技术文档中心检索
  • 为驱动器、控制器与运动系统交付生产级排障访问

Content Score

2023 · 维护中 · 负责人

内容质量评分——加速营销漏斗决策的 ML 信号。

  • 构建用于漏斗优化的内容质量评分后端分析平台
  • 开发 ML 工作流以评估内容效果并优先推进转化导向改动
  • 将评分输出与营销及内容工作流连接,加快数据驱动决策

Partner Locator

2022 · 进行中

CRM 到搜索数据集成——使合作伙伴与地点记录可靠可发现。

  • 设计 7+ 个 MuleSoft API 工作流,用于合作伙伴账户与地点的 CRM 增量更新
  • 将多源记录集成到 Lucidworks Solr 与 Elasticsearch 搜索,并优化缓存
  • 提升合作伙伴与地点数据在可搜索企业应用中的可靠性
  • 持续对生产级合作伙伴定位体验进行主动维护与增强

AEM 中的评估问卷——面向网络安全准备的第一方偏好采集。

  • 定制带分类支持、报表与仪表盘展示的 AEM 问卷组件
  • 支持生成基于 NIST 评分与定制报告的网络安全准备评估
  • 改进评估驱动内容工作流的结构化创作与多渠道发布

Site Feedback

2022 · 维护中

站内反馈闭环——将 Hotjar 信号接入 AEM 管理页面。

  • 将 Hotjar 与 Adobe Experience Manager 集成,采集站内反馈与行为信号
  • 将反馈采集连接到 AEM 页面,便于团队识别 UX 摩擦并优先修复
  • 建立早期反馈闭环,为后续个性化与内容评分工作提供依据

教育背景

门罗大学

工商管理硕士(MBA),兼职 · 在读

2025 - 至今

选修课程:战略营销与数据挖掘、软件系统设计、计算机网络、管理决策研究与统计、21 世纪组织行为与领导力、全球环境中的管理

史蒂文斯理工学院

数据科学硕士(MSc)

2019年9月 - 2021年12月

相关课程:统计方法、统计推断、高级优化方法、高级数据分析与机器学习、深度学习、自然语言处理、网络分析、数据库管理系统、Web 编程、数据结构与算法

广州大学

数学与应用数学学士(BSc)

2014年9月 - 2018年5月

相关课程:概率论与数理统计、运筹学、数值分析、高等代数、数学分析、实变函数、泛函分析、常微分方程、偏微分方程

奖项与荣誉

教务长奖学金

史蒂文斯理工学院

2019

获颁教务长奖学金(Provost’s Scholarship)。

CUMCM

广州大学

2015

全国大学生数学建模竞赛全国二等奖;省级一等奖。

课外活动

  • UBS Quant Hackathon — 2020
  • UBS Pitch Competition — 2020
  • Stevens HealthTech Hackathon — 2019

精选课程项目

比特币影响因素分析

史蒂文斯理工学院

2021 · Hoboken, NJ

面向推文情感与市场指标的 NLP 与深度学习。

  • 清洗并分析 2000 万+ 相关推文,以秒级粒度提取情感特征
  • 工程化 MACD/RSI 类指标,并对多种深度学习模型进行回测以评估组合收益

量化 AI 对岗位技能的影响

史蒂文斯理工学院

2021 · Hoboken, NJ

结合聚类与主题建模的职位描述 NLP。

  • 抓取并结构化 Indeed 职位描述文本;用 K-means 技能集聚类过滤噪声
  • 用主题建模与方面抽取神经网络分析技能趋势

电商推荐系统

史蒂文斯理工学院

2020 · Hoboken, NJ

基于用户、物品与交互数据的协同过滤。

  • 基于京东交互数据,用基于记忆与基于模型的协同过滤构建推荐引擎
  • 在 7,000 个用户-物品样本上评估,达到 77% Top-10 准确率

MyPlace Web 开发

史蒂文斯理工学院

2020 · Hoboken, NJ

面向家具与租房信息交换的全栈 Web 应用。

  • 带领四人团队构建 Node.js/Express 应用,含 MongoDB schema 与 RESTful CRUD API
  • 实现认证、搜索、评论与仪表盘功能,支持信息发布与交换

金融科技路演竞赛

史蒂文斯理工学院

2020 · Hoboken, NJ

基于多源公共数据的城市活力建模。

  • 用 Google POIs、Instagram、Zillow 与 BLS 数据构建活力指数,并用 Pandas 做特征工程
  • 微调 XGBoost 与 Keras 模型,预测美国城市繁荣趋势