Skip to main content

Abdullah Md Raihan Chy

Almost everything you did today produced data. Unlocking your phone, sending a message, buying groceries with a card, watching a video to the end or closing it after nine seconds. None of that was recorded because anyone cares about you personally. It was recorded because storage is cheap and someone might find it useful later.

Data science is the work of turning that pile into something a person can act on.

That is the whole idea. Everything else is method.

Why it exists now and not before

People have analysed numbers for centuries. What changed recently is scale.

Thirty years ago a company’s data fit in a filing cabinet or a modest database. An analyst could look at it. Today a mid-sized online shop generates more records in a week than that entire filing cabinet held, and a large one generates more in an hour.

This is what “big data” means, and the phrase is usually explained through three properties.

  • Volume. There is far too much of it to read.
  • Velocity. It arrives continuously rather than in batches you can sit down with.
  • Variety. It is not neat rows and columns. It is text, images, video, voice recordings, sensor readings and clicks, all mixed together.

Past a certain point, human reading stops being possible. You cannot look at ten million rows. You need methods that summarise, find patterns and make predictions on your behalf. That is where data science begins.

What the job actually involves

There is a gap between how this work is advertised and what it is like to do. Most descriptions emphasise building models. Most of the actual time goes elsewhere.

Getting the data

Before anything, you have to obtain it. Sometimes it sits in a company database. Sometimes it is spread across several systems that were never designed to talk to each other. Sometimes you collect it yourself with surveys, sensors or a program that reads public websites.

Cleaning it

This is the largest part of the job and nobody mentions it in the brochures.

Real data is a mess. Dates are written five different ways in the same column. Someone’s age is recorded as 250. Half the phone numbers have no country code. The same customer appears three times under slightly different spellings of their name. Entire fields are blank because the person filling the form skipped them.

All of that has to be found and dealt with before any analysis means anything. A model trained on dirty data produces confident nonsense, and it will not warn you. Experienced practitioners commonly say this stage takes most of a project, and my own experience matches that.

Exploring it

Once the data is usable, you look at it before you model it. Average values, ranges, how things relate to each other, simple charts. The point is to understand what you have and to notice anything strange early. Skipping this step is how people end up explaining a beautiful result that turns out to be a data entry error.

Modelling

Now you build something. This means fitting a mathematical model to the data so it can either explain what happened or predict what comes next.

Communicating

The final step, and the one that decides whether any of the work mattered. A result nobody understands changes nothing. Being able to explain a finding to someone with no technical background is not a soft extra skill in this field. It is the point of the field.

The three main toolsets

Statistics

The oldest and most important. Statistics tells you whether a pattern you found is real or whether it appeared by chance. This matters more than it sounds. Look at enough data and you will always find something that looks like a pattern. Statistics is the discipline that stops you believing it.

Machine learning

Machine learning means writing a program that learns rules from examples instead of being given the rules directly.

Here is the difference in plain terms. The traditional way to detect spam email is to write instructions by hand. Block anything containing certain words. Block anything from certain addresses. This works until spammers change their wording, which they do immediately.

The machine learning way is to show a program a hundred thousand emails already labelled as spam or not spam, and let it work out for itself what distinguishes them. Nobody writes the rules. The program derives them from the examples, and when spam changes you retrain it on newer examples rather than rewriting your logic.

This approach powers most of what people now call artificial intelligence. It is the reason your phone recognises faces in photographs, the reason streaming services suggest what to watch next and the reason banks can spot a stolen card in seconds.

Data visualisation

Turning numbers into pictures, because humans read pictures far better than tables. A chart that reveals a trend in two seconds is doing real analytical work, not decoration. Choosing the wrong chart, or a misleading scale, does the opposite just as effectively.

Where this shows up in practice

Healthcare. Models trained on medical scans can flag signs of disease that are easy to overlook, giving radiologists a second opinion. This is particularly valuable in countries where specialists are scarce and patients wait months for a reading. It also raises hard questions about responsibility when the model is wrong, and those questions are not solved.

Banking. Every card transaction is scored for risk in the moment it happens. The system compares it against your normal behaviour and decides in milliseconds whether to approve, verify or block. My own research is in this area, and the difficult part is not detecting fraud. It is detecting fraud without constantly blocking legitimate customers who happen to be travelling.

Agriculture. Photographs of crops, taken by phone or drone, can be analysed to identify disease early enough to treat it. For a country where a large share of the population depends on farming, catching a potato blight two weeks sooner is not a technical curiosity. It is a harvest.

Retail. Predicting what will sell and when, so shops hold the right stock. Unglamorous and enormously valuable, since both empty shelves and unsold inventory cost money.

Transport. Traffic prediction, route planning and estimating when a bus will actually arrive rather than when the timetable claims it will.

Where it is heading

Real-time analysis. Decisions made as data arrives rather than in a report next month. Fraud detection already works this way. Most things do not yet.

Wider access. Tools are becoming usable by people who are not programmers. This is genuinely good, and it carries an obvious risk. Software that produces a confident answer regardless of whether the underlying data supports it will produce a lot of confident wrong answers.

Explainability. Many powerful models cannot easily explain their reasoning. When such a model decides on a loan application or reads a medical scan, “the system said no” is not an acceptable answer. Building models that can justify their outputs is now a major research direction, and one I work in.

Privacy. Methods that allow models to learn from data without that data leaving the device it lives on. This matters more every year, for reasons that need no explanation.

If you want to start

I teach computer science, and this is the question students ask most often. The honest answer is shorter than they expect.

Learn Python. Learn basic statistics, properly, because it is the part people skip and the part that separates useful analysis from confident nonsense. Then find a dataset about something you actually care about, cricket scores, your own expenses, local weather, and try to answer one real question with it.

You will spend most of your time cleaning the data and being frustrated. That is not a sign you are doing it wrong. That is the job. Everything else is easier than it looks, and that part is harder than it looks.


I research machine learning, quantum computing and AI-driven cybersecurity, and teach Computer Science at Sunshine Grammar School and College in Chattogram. My papers are on the research page.

4 Responses

  1. I made the switch a month ago and the transition has been completely seamless. The customer team responds in a heartbeat and the daily bonuses actually boost my bankroll noticeably. It is refreshing to play at a place that truly puts us first. 776betapp

  2. I love how lightweight their app feels on my device. No lag at all even during peak hours when the big matches are on. The daily promotions keep things fresh and I never feel stuck waiting for a payout. It is rare to find a betting site that balances speed and fun this well. batterybetgame

  3. I’ve tried a few sites lately but this one really stands out. The interface is smooth and the payouts are surprisingly fast. Definitely checking out bdgv88 again tonight!

Leave a Reply

Your email address will not be published. Required fields are marked *