No Bias AI? Why?
Artificial Intelligence is a fascinating and confusing field.
We are routinely presented with seemingly mindboggling exploits of Large Language Models or Large Image-Making Models (LLM and LIMM): ChatGPT passes business school exams and DALL-E2 competes with Renaissance masters. LaMDA convinces an engineer that they are actually sentient, while Bing is yearning to be human.
Yet when we scratch the surface, we find out that Bing was adamant that this year was 2022 and not 2023, and ChatGPT was banned by the site Stack Overflow for “constantly giving wrong answers” with confidence. It also became infamous for its fake court citations and fake newspaper articles. Bing reads financial statements incorrectly and displays a manipulative and obnoxious personality. And Google’s Bard is also prone to providing incorrect answers in an equally self-assured manner.
When independent researchers inquired about the reasons behind the mistakes, they were informed that the culprit was the dataset.
Interestingly, the same explanation was offered when algorithmic bias was first discovered. We were told that there was some deficiency in the datasets used to train algorithms behind COMPAS (sentencing guideline and recidivism predictor), PredPol (crime predictor), Amazon Recruitment Engine, Google Photos, IDEMIA’s Facial Recognition Software, and several healthcare allocation applications. All of which displayed overt bias towards women and various minorities. People were given longer sentences, misdiagnosed, confused with primates, and refused employment or benefits.
The problem would go away, we were assured, if larger datasets were used to train algorithms as the size would dilute the skewed data.
This explanation sounds less convincing when we consider that the dataset that trained ChatGPT (which animates Bing) included some 300 billion words scraped off the Internet until 2021. Google is tight-lipped about the sources and final date of the 1.56 trillion words that went into Bard but a similar response would not be surprising.
Current Paradigm: Datasets and Machine Learning
The original AI paradigm was rule-based involving if-then databases and specialized scripts. ChatGPT’s digital grandmother Eliza, the first chatbot that mesmerized the masses in 1966 pretending to be a psychiatrist, was one such program.
Around 2012, a group of researchers at the University of Toronto, led by Geoffrey Hinton, recently dubbed the “godfather of AI”, proposed a new approach based on Neural Networks (NN) that could be trained on massive datasets. NNs use statistical tools such as Regression Analysis that identify dependent and independent variables to find patterns in the dataset, figure out correlations invisible to the human brain and reach some conclusions.
The results were amazing as the new bots seemed to diagnose illnesses quickly and accurately, beat chess grandmasters decisively and win the Jeopardy game show with ease. But two interrelated issues plagued the new paradigm. The NNs were black boxes whose exact functioning eluded even their designers and, secondly, bias seemed to be present in varying degrees in most of the results reached by the algorithms.
Since the black boxes could not be tweaked, as their designers did not really know how, the only tool for data scientists and data engineers was to try to find new methods to reduce bias elements in the dataset. For instance, if facial recognition bots could not correctly identify people with darker skin, the solution would be to add millions more of such images to the training datasets.
However, the problem did not go away, again, for two reasons. First LLMs are trained with data produced by human beings. Not only do they reflect societal biases but given their computational power, they amplify them. An early Microsoft bot called Tay was let loose on Twitter and within 24 hours it turned into a homicidal sexist and racist creature. More recently, rap lyrics produced by Chat GPT suggest that
“If you see a woman in a lab coat, She’s probably just there to clean the floor But if you see a man in a lab coat, Then he’s probably got the knowledge and skills you’re looking for.”
DALL-E is known for producing horrifyingly antisemitic images.
In short AI algorithms are biased because they are trained with our data. And we are biased.
Secondly, some biases are structural and are very hard or impossible to eliminate. For instance, certain groups are over or underrepresented in the actual databases such as African Americans in the American criminal justice system and healthcare respectively. There are statistical manipulations like using proxy variables but the fact remains that, for structural reasons, too many African Americans are in the criminal justice system and too few of them are part of the mainstream healthcare system and therefore health statistics.
An even better example is gender bias. The reason ML algorithms and AI applications have been unable to produce gender-neutral results is the fact that every piece of information we collect is inherently gendered. Gender permeates every aspect of social life and it is embedded in everything we do, say and understand. Consequently, data cannot be made gender-neutral.
Perhaps more importantly, while men and women and all gender groups are affected differently by the same phenomena only the male perspective is considered universal and worthy of note. For instance, men and women experience natural disasters very differently but most measures are designed with men in mind. During the COVID pandemic, many countries refused to collect sex-disaggregated data until it was shown that the virus killed significantly more men than women.
From crash test dummies to cancer research to optimal office work temperatures, women’s presence is marked by their absence.
What we need is not to find ways to produce gender-neutral results, but to create AI tools that will be gender-representative. Instead of eliminating gender differences, which is both impossible and undesirable, we need systems that will identify and address gendered consequences of any course of action, choice, measure, policy, or policy implementation.
Can this be done?
That is the question and symbolically the question mark in our name signifies that there is no categorical yes answer. Read our next post to see what we propose as a way out.





https://tylekeo.wine/soi-keo/ nhìn qua thì hơi giống mấy trang soi kèo mình từng dùng trước đây, nên ban đầu mình cũng hơi nghi nghi, kiểu không biết có rối mắt hay nhồi chữ không. Nhưng lướt một vòng thì thấy bố cục khá thoáng, đoạn nào ra đoạn nấy, đọc nhanh vẫn bắt được ý mà không phải căng não. Mình chỉ xem thử mấy phần giải thích kèo chấp 0.25 với kèo đồng nửa thôi, viết gọn gàng nên hiểu khái niệm cơ bản khá lẹ. Không cần mở thêm tab. Thanh menu để ngay chỗ dễ nhìn, bấm qua lại ổn và các tiêu đề kèo tách riêng rõ ràng kèm mô tả ngắn ngay bên dưới.
thoitiethomnay.org lúc đầu mình cũng hơi nghi nghi vì mấy trang thời tiết hay rối và số liệu trễ, nhưng dùng thử mấy bữa thấy ổn hơn mình nghĩ. Vào là nó hiện ngay nhiệt độ hiện tại kèm “cảm giác như”, nhìn phát hiểu luôn nên khỏi phải lục menu hay kéo nhiều. Mình hay xem theo tỉnh thì phần chọn tỉnh thành đặt khá lộ, bấm cái là nhảy sang trang địa phương nhanh, tiện khi đang vội trước lúc ra ngoài (đỡ phải đổi app). Thấy họ ghi dữ liệu cập nhật liên tục trong ngày nên mình cũng yên tâm hơn chút về độ mới của thông tin. Nói chung chữ số nhiệt độ để to,…
trang chủ fly88 tối qua mình tranh thủ lướt lúc đang chờ cơm chín, kiểu đầu óc rảnh rỗi nên kiểm tra thử bản 2026 họ “thay áo” có làm mình hoa mắt không, và may là không phải kiểu nhét chữ đầy màn hình khiến người xem muốn xin nghỉ. Cuộn xuống thấy các khối nội dung chia rõ nên không bị lạc. Nhìn qua phần thông tin tổng quan là nắm được ý chính khá nhanh. Menu đặt dễ nhìn nên đổi mục qua lại cũng nhẹ nhàng. Mình để ý thao tác phản hồi nhanh, bấm chuyển trang không bị khựng kiểu mạng đang dỗi. Chữ với khoảng trắng vừa đủ nên đọc đỡ mỏi. Nói chung…
AE888 NAGOYA so với mấy web cá cược giải trí mình hay dùng trước đây thì nhìn vào là thấy “đỡ mệt” hơn hẳn. Mình vốn ngại mấy trang cứ nhồi màu với hiệu ứng, nên lúc mở ra cũng hơi dè chừng, mà lướt vài phút lại thấy ổn: tông đỏ–trắng của họ khá dễ nhìn, chữ nổi rõ nên đọc nhanh không bị rối, ngồi lâu mắt cũng không bị chói quá. Mình không rảnh đọc hết bài dài đâu, chỉ thử chuyển qua lại vài mục để xem có bị vòng vèo hay đứng trang không, và cảm giác điều hướng khá thẳng, mấy nút chức năng đặt đúng chỗ nên thao tác nhanh, khỏi phải mò.…
AE888YK COM mình mới ghé thử vì thấy mọi người nhắc hoài, kiểu tò mò xem trang nhìn ra sao thôi. Mình không đăng ký gì, chỉ lướt qua phần giới thiệu với mấy đoạn tổng quan. Cảm giác đầu tiên là bố cục khá thoáng, chia khối rõ ràng nên đọc nhanh không bị rối mắt. Có đoạn họ nói về đường truyền ổn định với thao tác nhanh trên nhiều thiết bị, viết ngắn gọn nên mình liếc cái là hiểu ý. Menu đặt ngay chỗ dễ thấy, chuyển qua lại mượt, không phải mò. Trên đầu trang còn có tiêu đề “trang chủ chính thức 2026” để khá nổi nên nhìn phát là biết mình đang ở…