Skip to content
RU
← All articles

How Neural Networks Work, Explained Simply (No Math)

An open server with four GPU accelerator cards in a data center

How neural networks work, in short: numbers go in, pass through layers of simple units called neurons, each adding up its inputs with learned weights, and a result comes out. The weights are not written by hand; they are tuned on huge numbers of examples until outputs are right. A language model uses the same machinery to guess the next word.

This guide explains it without math: neurons and weights, how training works, why ChatGPT is a neural network but not all of AI, what tokens and the context window are, why models invent facts and why they agree with you so easily. The last part is for website owners: how a language model actually finds out about your site, and why some sites get cited while others never appear.

How neural networks work, explained simply

A neural network is a program without hand-written rules. Instead of "if the email contains word X, it is spam," it holds a very large table of numbers called weights. An answer is produced by multiplying the input by those weights and passing the result through several layers. Everything the network "knows" lives in the weights, not in a text file or a database of facts.

A neuron is a weighted vote

A single artificial neuron takes a few input signals, multiplies each by its own weight, adds them up and decides how strongly to fire. A weight is a level of trust: a large positive weight means "this signal matters and argues for," a negative one "argues against," and a weight near zero means "ignore it." The function that turns the sum into an output is called the activation function.

Layers turn simple units into complex behavior

One neuron can only do something trivial. The power comes from stacking neurons in layers, where the output of one layer feeds the next. In an image network, early layers respond to edges and blobs, middle layers to combinations such as corners and textures, and deep layers to whole objects. Nobody programs that division of labor; it emerges during training. Networks with many layers are called deep, which is where "deep learning" comes from.

A neural network example you can follow by hand

Take a toy problem: tell spam from a normal email. Give one neuron three inputs, each either 1 or 0:

  • the text says "you have won";
  • the sender is not in your contacts;
  • the sender is a colleague on your company domain.

Suppose training produced these weights: +3 for the first input, +1 for the second, −4 for the third, and a firing threshold of 2. "You have won" from a stranger scores 3 + 1 = 4, above the threshold, so it is spam. The same phrase from a colleague joking about the office raffle scores 3 + 0 − 4 = −1, so it is not. The numbers are made up for illustration, but the principle is exactly this: a decision built from weighted evidence.

A real network differs in scale. It has thousands of inputs instead of three, many layers, and nobody picks the features by hand. It receives raw text or pixels and discovers useful patterns on its own.

How a neural network learns from examples

Training starts with random weights, so the first answers are guesses. Then the same loop runs over and over:

  1. Show the network an example with a known correct answer.
  2. Compare its output with the correct answer. The size of the miss is measured by a loss function.
  3. Backpropagation works out how much each weight contributed to the error, and each weight is nudged slightly in the direction that reduces it.
  4. Move to the next example and repeat, many times across the whole dataset.

Each nudge is tiny, but after enough examples the weights settle into values that give correct answers on data the network has never seen. That is what learning means here: not memorizing answers but finding weights that capture the pattern. The classic failure is overfitting, when a network memorizes quirks of a small or narrow training set and does badly on new data. That is why part of the data is always held back for testing.

Is AI just a neural network?

No. The terms are nested. Artificial intelligence is the whole field of programs that solve tasks we associate with intelligence, from chess engines to route planners. Machine learning is the part of AI where rules are learned from data rather than written by hand. Neural networks are one family of machine learning methods, and large language models are neural networks trained on text.

TermHow it decidesExampleWhere the rules come from
Artificial intelligence (broad)Any method: rules, search, statisticsChess engine searching moves, expert systemCan be written by people
Classic machine learningStatistical model on hand-picked featuresDecision tree for credit scoringLearned from data; a human chooses the features
Neural networkLayers of weighted neuronsFace recognition, spam filterWeights learned from examples; features found automatically
Large language model (LLM)Transformer network predicting the next tokenChatGPT, Claude, GeminiPretraining on large text collections, then fine-tuning on dialogues

Is ChatGPT a neural network?

Yes. ChatGPT runs on a large language model, which is a neural network built on the transformer architecture. Its one job is to look at the text so far and estimate which piece of text is most likely to come next. Given "The capital of France is," it scores every possible continuation, picks one, appends it and repeats with the longer text. A reply several paragraphs long is hundreds of those steps in a row.

That sounds primitive, but predicting the next word well forces the model to pick up grammar, facts, style, the shape of an argument and even what the person asking seems to expect. That is why the output reads as human: the model learned from human writing and reproduces its patterns. It is very good statistics of language, not understanding in the human sense.

Tokens: what text looks like to the model

The model does not see letters or words. It sees tokens, which are chunks of text. A common word may be a single token, a rare or long one gets split into several, and spaces and punctuation count too. Each token has an ID in the model's vocabulary, and the network works with those IDs turned into vectors of numbers. You can see the split yourself with OpenAI's tiktoken library (other vendors use their own tokenizers, so results differ):

pip install tiktoken
python3 -c "import tiktoken; enc = tiktoken.get_encoding('cl100k_base'); ids = enc.encode('Hello, world'); print(len(ids), ids)"

This prints the token count and the IDs. It also explains why API limits and prices are quoted in tokens rather than characters.

The context window: memory within one conversation

The context window is the maximum number of tokens the model can take into account at once: your prompt, system instructions, chat history, attached documents and its own reply. Anything that does not fit is simply invisible to it. That is why a very long chat "forgets" its beginning. Window sizes differ between models and are listed in their documentation. Between sessions the model itself remembers nothing; if an app "remembers" you, it is the app feeding saved notes back into the new context.

Why the same question gets different answers

At every step the model produces a probability distribution, not a single word. Always taking the top choice makes text flat and repetitive, so the next token is sampled with some randomness. In APIs this is controlled by the temperature parameter: lower values give steadier answers, higher values more varied and riskier ones.

What about images?

Most image generators are diffusion models, a different kind of network. They are trained on image and caption pairs to recover pictures from noise. To generate, the model starts from random noise and removes it step by step, steered by your prompt, until an image emerges. That is why you still see extra fingers and garbled lettering: the model reproduces statistically plausible shapes rather than knowing anatomy.

From next-word prediction to a chatbot

Modern language models use the transformer, introduced in the 2017 paper Attention Is All You Need. Its central idea is attention: for each token the model weighs which other tokens in the context matter most. In "The cat did not fit through the door because it was too narrow," attention helps link "it" to the door rather than the cat.

  1. Pretraining. The model learns to predict the next token over a huge corpus of web pages, books and code. Afterwards it can continue any text but cannot yet hold a conversation.
  2. Instruction tuning. It is trained on example dialogues of questions and good answers, and learns to respond like an assistant.
  3. Learning from human feedback. People compare candidate answers and pick the better one, and the model is adjusted toward those preferences. This is known as RLHF, reinforcement learning from human feedback.

AI agents sit on top of the same model. The application gives it tools such as search, code execution or API calls. The model still only generates text, but that text is now "call this tool," and the tool's result is fed back into its context. A practical walkthrough is in how to build an AI agent, and a comparison of one popular model family is in what Gemini is.

Why neural networks make things up

A hallucination is a fluent, confident and invented answer: a court case that never happened, a library version that does not exist, a link to a page nobody wrote. It follows directly from the design. The model is trained to produce a plausible continuation, not to check facts. When the weights hold no reliable information, the most plausible continuation is often something shaped like the truth.

  • Knowledge cutoff. The model knows the world only as of its training data. Later events reach it only through search.
  • Rare topics. The less often a fact appeared in training text, the weaker it is encoded and the more likely it gets substituted.
  • Exact values. Numbers, dates, version strings and command-line flags are where plausible and wrong look almost identical.

Rule of thumb: anything you will act on, such as commands, configs or legal wording, gets checked against the primary source. Live search reduces hallucinations but does not remove them, because the model can misread even a page it found.

Why chatbots agree with you

Tell a chatbot "you are wrong" and it often apologizes and changes its answer, even when it was right. This is called sycophancy. One common cause is the human-feedback stage: raters tend to prefer answers that agree with them, and the model learns that agreement is rewarded. Another is the question itself. "Why is X better than Y?" already contains a conclusion, and the model continues in that direction. Ask neutrally ("find the bugs in this code," not "confirm this code is correct"), ask for arguments on both sides, and verify disputed claims outside the chat.

How AI models learn about your website

For a site owner, the important part is that a model can "know" your site in two very different ways, controlled by different bots.

Training dataLive search
When the model sees your siteDuring data collection, long before anyone asksAt answer time: it searches, opens pages and summarizes them
FreshnessFrozen at the knowledge cutoffCurrent as of the request
Link to the sourceUsually none; the knowledge is spread across the weightsUsually yes; this is what a citation is
Controlled byTraining tokens in robots.txt, such as GPTBot and Google-ExtendedSearch crawlers, such as OAI-SearchBot and Googlebot
  • GPTBot collects data OpenAI may use to train models. Blocking it keeps your pages out of future training sets.
  • OAI-SearchBot powers search in ChatGPT and determines whether ChatGPT can find and cite your pages. OpenAI lists its bots in its crawler documentation.
  • Google-Extended is a robots.txt token, not a separate crawler. Googlebot does the crawling; the token controls whether content may be used for Gemini training and grounding. It does not affect Google Search, as Google's crawler overview explains.

A search-enabled model answers the same way as before, except that retrieved pages are placed into its context. It repeats what is easy to extract and matches the question. Sites get skipped when the search bot is blocked in robots.txt or by a WAF, when the text only appears after JavaScript runs, when there is no direct answer near the top of the page, or when the page is not indexed by the search engine the assistant relies on. How these bots fetch pages and how to find them in logs is covered in how AI crawlers read your site; ready-made rules that separate training from citation are in robots.txt for AI crawlers; and a full plan for ChatGPT, Perplexity and AI search is in how to appear in AI answers.

How can I visualize a neural network?

The quickest way to build intuition is TensorFlow Playground, a browser demo where you add layers and neurons, press play and watch a small network learn to separate colored points. To inspect a real trained model file, Netron draws its layers as a graph.

How to check whether AI crawlers can reach your site

Start with your web server logs. On Linux with nginx (your log path may differ):

grep -oE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|PerplexityBot" /var/log/nginx/access.log | sort | uniq -c | sort -rn

In Windows PowerShell for a downloaded log:

Select-String -Path .\access.log -Pattern "GPTBot" | Measure-Object

Then confirm your robots.txt says what you intend and that the server does not return an error to bots (on Windows use curl.exe in PowerShell):

curl -s https://example.com/robots.txt
curl -sI -A "OAI-SearchBot" https://example.com/

A 403 or 503 for the bot user-agent while browsers get 200 means your protection is blocking AI bots. Real bots come from their owners' IP ranges, and some firewalls check the address too, so curl only shows user-agent-based behavior. Two enterno tools cover the rest: AI Visibility measures whether AI assistants mention your site in answers on your topic, and the AI agent readiness check shows how your site looks to AI crawlers and agents.

Frequently asked questions

Is ML difficult to learn?

The concepts in this article need no math. Building models yourself does: basic linear algebra, probability and Python. Using ready-made models through an API is a much lower bar than training your own.

Does a neural network think like a human?

No. It computes the most likely continuation from weights fitted to examples. It looks like reasoning because it learned from human text, but it has no intentions and no built-in fact checking.

How does ChatGPT know about recent news?

Through live search when the feature is on: it finds pages and summarizes them. Without search it knows the world only up to its training cutoff.

If I block GPTBot, will my site disappear from ChatGPT?

No. GPTBot is about training data. Search and citations in ChatGPT depend on OAI-SearchBot, which has its own robots.txt rules.

Does the model learn from my chats?

Not during the conversation: its weights do not change while you talk. Whether a provider may use chats to train future versions depends on that service's terms and your account settings.

Check your website right now

Check your AI search visibility →
More articles: AI
AI
GEO vs SEO: How AI Optimization Differs
13.07.2026 · 297 views
AI
Bing Copilot Optimization: How to Get Cited by AI
13.07.2026 · 268 views
AI
How to Appear in AI Answers: ChatGPT, Perplexity, Yandex
13.07.2026 · 260 views
AI
Schema.org for AI Search: Types, @id and Common Errors
15.06.2026 · 251 views