Data Science team, with three years across search, ML pipelines, file sync, chatbots, and network analytics, applying ML to network-based problems.
Accomplishments
- Led file-sync tool development: a delta-sync service in C++ substantially reduced sync-server load; also built a stateless updater, a multi-tenant client update framework, and a rate limiter that improved sync frequency 30%.
- Built a financial-document search engine with +45% search relevancy, +30% online ontology enhancer, with an NLP pipeline for clustering, keyword extraction, and text classification.
- Designed an ML model object-storage / delivery pipeline (gRPC, protobuf, Redis, RabbitMQ, S3): +40% deployment frequency, clients in Java / Python / Go / C++, weekly release cadence.
- BLE indoor location: 95% region classification for static assets, 68% regression accuracy via RSSI triangulation + interference correction; a live asset-view portal added 70% correctness to site calibration.
Full detail Hide detail
More Talentica work
- Network traffic estimation: an ensembled regression model for bandwidth prediction from SLA metrics and native Speedtest feedback.
- Network traffic identification: a multi-modal ML model (regression ensemble + auto-encoders) classifying audio vs. video streams.
- Real-time ML model deployment pipeline on Storm, Kafka, Python, Java.
- AWS Athena data-dump tool: an ad-hoc Storm-topology client feeding a disk + batch pipeline (Java, Spark).
- Single-cell identity classification: an auto-encoder network, 83% accuracy for cell-based print technology.
- Data platforms
- Machine learning
- NLP / search
- Distributed pipelines