Big Data Development Engineer Intern
About the role
About Us
Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance. Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution. Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services. Key Responsibilities:- Research & Benchmarking — Survey industry and open-source approaches in AI + data governance (metadata & lineage platforms, semantic layers, data quality frameworks, NL2SQL / NL2Metric); produce a best-practice proposal adapted to our tech stack.
- Data Model Governance — Build AI-assisted model review capabilities: naming convention and layering (ODS/DWD/DWS/ADS) validation, duplicate and redundant model detection, lineage-based identification of unused / low-value / high-cost assets, and automated refactoring recommendations.
- Metric Semantic Automation — Automatically extract and standardize metric definitions from SQL, lineage, and documentation; detect definition conflicts and redundant builds of the same metric across different reports; maintain a machine-readable semantic layer to support natural language metric queries.
- Data Quality Automation — Auto-generate quality rules based on data profiling and lineage (replacing hand-written rules); perform anomaly detection on data volume, distribution, timeliness, and schema drift; conduct root cause analysis along lineage; implement automated alert grading and remediation recommendations (or auto-remediation).
- MVP Delivery — Run at least one end-to-end implementation across the three areas: problem definition → solution design → prototype → deployment on a real data domain → quantified results (coverage, precision/recall of issue detection, manual effort saved) → iteration.
- Documentation & Communication — Produce solution designs, evaluation methodologies, and results; present findings to platform and data stakeholders; deliver a reusable framework rather than one-off scripts.
- Undergraduate or graduate student in Computer Science, Data Science, Statistics, or a related field.
- Solid proficiency in SQL and Python. Understanding of data warehouse fundamentals — dimensional modeling, layered architecture, metadata, and lineage. Experience with Spark / Flink / Hive / StarRocks is a plus.
- Hands-on experience with LLM application development: prompt engineering, RAG, Agent / tool-calling frameworks (e.g., LangChain, LlamaIndex, MCP), with the ability to evaluate whether an LLM system is actually effective.
- Structured thinking: able to distill a vague governance pain point into a well-defined problem with quantifiable success criteria, and honestly articulate what the MVP validated and what it did not.
- Self-driven and comfortable with ambiguity — this is an exploratory project with no predetermined answers.
- Bonus: experience with DataHub / OpenMetadata / Atlas, dbt, Great Expectations / Deequ, or any metrics / semantic layer tooling.
- Able to read technical materials in English; clear written communication skills.
- Minimum 3-month internship commitment, 5 days per week on-site.
- Fluent in Mandarin is required; fluent English is a plus.
Why Join Us
At Bybit, we are committed to fostering a supportive and enriching work environment.
Our benefits include:
- Study Growth Fund: We support your professional development and continuous learning.
- Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation.
- Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world.
- Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company.
- Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.