OROVA.VN — BIZ AI AGENT
Guides

ChatGPT for SEO: Nine Jobs It Does, Four It Ruins

Orova 7 views
ChatGPT for SEO: Nine Jobs It Does, Four It Ruins

You have the tab open next to your keyword sheet, and it has already saved you an afternoon this month. It has also, at some point, handed you a search volume figure that was invented, a citation to a study that does not exist, and a paragraph so smooth that nobody on the team noticed it said nothing. Both of those things are true at once, and that is the whole problem with the advice available on this subject.

Most guides to ChatGPT for SEO are prompt lists. Prompt lists are the least useful half of the answer, because a prompt only matters once you have decided the job is a good fit. The useful half is the boundary: which jobs a language model can finish, which ones it can only start, and which ones it will complete confidently and wrongly — the dangerous category, because the failure leaves no visible mark on the output. A hallucinated statistic looks exactly like a real one.

So this piece draws the line job by job. Nine tasks you can hand over today and roughly what you have to supply for each. Four where handing over breaks something, with the reason each one breaks rather than a general warning. The prompt shape that returns work you can publish instead of work you have to rewrite. The tells an editor uses to spot machine prose in thirty seconds. What Google has actually written down about AI-generated content, quoted rather than paraphrased, because the folklore version is wrong in both directions. And the point at which a chat window stops being enough, which arrives sooner than most teams expect.

Which SEO jobs can you hand to ChatGPT?

Hand over the jobs where you own the facts and the model supplies the shape: clustering a keyword list, drafting outlines, writing titles and meta descriptions, generating schema, summarising a page, rewriting for clarity, drafting FAQ questions, translating a first pass, and checking tags for consistency. Keep the jobs that require data the model cannot see.

That sentence is the whole framework, and it is worth stating in its general form because it survives model upgrades. A language model is very good at transforming text it has been given and very unreliable at retrieving facts it has not. Every safe job on the list is a transformation. Every unsafe job is a retrieval wearing a transformation's clothes.

Hold that distinction and you can classify a new task in about five seconds, which is more useful than any prompt anyone will sell you.

The nine jobs, and what you have to supply for each

One: clustering a keyword list. Paste four hundred keywords with their volumes and ask for groups by search intent, one group per page. The model is good at this because the grouping is a text-similarity judgement, which is exactly what it is built on. What you supply is the list and the volumes — it must not invent either. What you check is that the clusters map to pages you would actually write, not to abstract themes.

Two: drafting an outline from a brief. Give it the target query, the audience, what the top results currently cover, and what you intend to say that they do not. It returns a structure faster than you can type one. The value is not the structure it invents; it is that a structure on the page is easier to argue with than a blank one.

Three: writing titles and meta descriptions. Twenty variants in one pass, then you pick. Supply the character limits and the exact query, and check every one of them against the page, because it will happily promise something the page does not deliver. This is the highest-value-per-minute job on the list and almost nobody does it in bulk.

Four: generating structured data markup. JSON-LD for an article, an FAQ, a product, a how-to. It knows the shapes. Two conditions, both from Google's own published guidelines: the markup must be a true representation of the page content, and you must not mark up content that is not visible to readers. Models will happily generate a FAQ block for questions that do not appear on the page, which breaks both rules at once. Generate, then validate, then check the page actually contains what the markup claims. If schema is not yet second nature, the plain-English version of structured data is worth twenty minutes before you start generating any.

Five: summarising the current results for a query. Paste in the visible text of the top pages and ask what each one covers and what none of them covers. You are asking it to compare documents you supplied, which is safe. Asking it what ranks for a query is not, because it cannot see the results page — that is job number two on the ruined list.

Six: rewriting a paragraph for clarity. Give it your paragraph and a specific instruction: shorter, plainer, less hedging, address the reader directly. Generic instructions like "make it better" return generic prose. Specific instructions return specific edits. This is the single job where the output most often goes straight in unchanged.

Seven: a first-pass translation. Good enough to review, never good enough to publish unread. The failure mode is not grammar — it is idiom that is technically correct and locally wrong, and a native speaker catches it in one pass. Use it to cut translation cost by two thirds, not to remove the reviewer.

Eight: drafting FAQ questions. Feed it the article and ask what a reader still would not know. It is good at spotting the unanswered question because it is essentially doing coverage analysis on text you gave it. Then answer them yourself, or it will answer them with plausible filler.

Nine: consistency checks across tags. Paste fifty title tags and ask which ones break your pattern, duplicate each other, or exceed the limit. It is a fast, tolerant pattern-matcher over a list you supplied. It will miscount characters, so verify anything near the limit with an actual character count.

Table of nine SEO tasks that can be delegated to a language model, each with what the human must supply, what the model returns, and the one thing to check before the output is used
The middle column is the one that decides. If you cannot supply it, the job belongs on the other list.

Notice what every one of those nine has in common. You brought the facts. The model brought the arrangement. Nothing on the list requires it to know something about the world that you did not put in front of it.

The four it ruins, and why each one breaks

These are not "be careful" warnings. Each one fails for a specific mechanical reason, and knowing the reason is what stops you making the same mistake with next year's model.

One: search volume and keyword difficulty. Ask for the monthly volume of a keyword and you will get a number. The number will be shaped correctly — round, plausible, in the right order of magnitude — and it will be invented. The model has no connection to any keyword database; it is producing the most likely next token given a question that looks like it should have a numeric answer. This is the most common and most damaging misuse in the whole category, because the output is indistinguishable from a real figure and it goes into a plan that people then spend money against. Volume comes from a keyword tool. If you want a refresher on where the real numbers come from and what they mean, the practical version of keyword research covers it without the vendor pitch.

Two: what currently ranks. Ask it who is in the top ten for a query and it will name plausible sites. Unless it has genuinely fetched the results page in that session, this is recall of a training set that is months old at best, mixed with reasonable guessing. Ranking is a live, personalised, location-dependent fact. It is not the kind of thing a language model holds.

Three: citations and sources. This one is worse than it looks, because the failure is partial. Ask for research supporting a claim and you may get a real institution, a real author, a plausible year, and a title that was never published — or a real paper whose findings do not say what the model says they say. Half-right is more dangerous than wrong, because half-right passes a skim. The only safe rule: if you have not opened the source yourself, it does not go in the article. There is no prompt that fixes this.

Four: judging intent in place of evidence. Ask what people mean when they search a phrase and you get a confident taxonomy. Sometimes it is right. It is a guess derived from language patterns, not from watching what searchers actually clicked, and it is most confidently wrong exactly where it matters most: commercially ambiguous queries where the money is. Use it to generate hypotheses about intent. Confirm them by looking at the real results page, which is the only place the answer lives.

Four panels showing the failure modes: invented search volumes, stale ranking claims, half-correct citations, and intent guessed rather than observed, each with the mechanism that causes it and the check that catches it
All four share one shape: the model was asked to retrieve rather than transform, and retrieval failures come out looking exactly like successes.

There is a fifth candidate people argue about, which is asking the model to evaluate its own output. It is not on the list because it is not exactly a failure — it is just weaker than people expect. A model will find real problems in text you give it, including text it wrote. What it will not do reliably is notice that a fact is false, because noticing that requires the knowledge it did not have when it wrote the sentence.

What Google has actually written about AI content

Two myths circulate, both wrong, both expensive. The first says AI-written content is penalised. The second says Google cannot tell and therefore it does not matter. Neither survives contact with the published documentation.

Screenshot of the Scaled content abuse section of Google Search Central's spam policies page, defining the practice as generating many pages for the primary purpose of manipulating rankings, with generative AI tools named in the first example
The policy is about purpose and value, not about which keyboard the words came from. That distinction does all the work.

The helpful content documentation says it plainly: if you use automation, including AI generation, to produce content for the primary purpose of manipulating search rankings, that is a violation of the spam policies. Read the sentence carefully, because every clause is load-bearing. The violation is defined by purpose, not by method. Automation is named explicitly, which kills the "they cannot tell so it does not matter" position — the policy does not depend on detection, it depends on what the page is for.

The spam policy page then defines the specific offence. Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users, and it is typically focused on creating large amounts of unoriginal content that provides little to no value to users, no matter how it is created. The first example given is using generative AI tools to generate many pages without adding value for users.

"No matter how it is created" is the phrase to remember. It cuts both ways. It means a thousand AI-written pages of nothing are a policy violation. It also means a thousand human-written pages of nothing are the same violation, which is a fact the industry produced at scale for fifteen years before any of this existed.

One more line worth knowing, because it changes what you do rather than what you believe. Among the self-assessment questions Google publishes is whether the use of automation, including AI generation, is self-evident to visitors through disclosures or in other ways. That is not a rule; it is a question they suggest you ask yourself. Teams that take it seriously tend to arrive at the same place: a named human author who actually reviewed the piece, and no pretence about the process. The trust component of E-E-A-T is the one Google calls most important, and an unreviewed byline is the cheapest way to lose it. If you want the fuller treatment of what that framework asks for in practice, the working version of E-E-A-T is the place to go.

The practical summary: the policy has nothing to say about your process and a great deal to say about whether the page was worth publishing. Which is, awkwardly, the same standard that always applied.

The prompt shape that returns publishable work

Most prompts fail for the same reason: they ask for an outcome without supplying the constraints that would let anything but a generic answer be produced. Six parts, in this order, and the difference between including them and not is roughly the difference between a draft and a rewrite.

Role and audience. Not "you are a world-class SEO expert" — that does nothing. Name the reader: someone who manages a small ecommerce site, has used Search Console twice, and needs to decide something this week.

The source material. Paste it. The single largest quality jump available in any prompt is the difference between describing your material and including it. Everything the model does not receive, it makes up.

The job, stated as a verb. Rewrite. Cluster. Compare. Extract. Not "help me with". A vague verb returns a vague artefact, every time.

The constraints, as numbers. Under sixty characters. Exactly eight bullets. No adjective before a noun unless it changes the meaning. Numeric constraints are followed far more reliably than adjectival ones, and they make the output checkable rather than arguable.

The forbidden list. The part almost everyone omits, and the part that removes most of the rewriting. No invented statistics. No sources you have not been given. Do not claim anything about rankings. Say "I do not have that" rather than estimating. Ban the words your industry has worn out.

The example. One paragraph of your own writing that sounds right. This does more for voice than any amount of describing it. "Professional but friendly" means nothing to a model and to be fair it means very little to a person either.

Diagram of a six-part prompt structure showing role and audience, source material, the job as a verb, numeric constraints, a forbidden list, and one example of the target voice, with what goes wrong when each part is missing
Read the right-hand column as a diagnosis. When output disappoints, one of these six is usually missing.

Then the part nobody mentions: give feedback on what came back and ask for a second pass. The second output is reliably better than the first, and the reason is not mysterious. Your correction is a constraint you failed to state up front. If you find yourself giving the same correction repeatedly, promote it into the prompt permanently. Over a few weeks this converges into something that is no longer a prompt at all — it is a specification, and it is the most valuable artefact your team will build in this area.

Six tells that give machine prose away

Useful whether you are editing your own output, auditing a freelancer, or working out why a page reads as thin without any single sentence being wrong.

Every paragraph is the same length. Human writing is lumpy. A one-sentence paragraph after a long one is emphasis. Uniform blocks of three to four sentences are the strongest single signal, and it survives every rewrite that does not address structure.

Balance where there should be a position. "While X has advantages, it also has disadvantages." A model reaches for symmetry by default because symmetry is safe. An expert has an opinion and says which one they would choose.

Specifics that are not specific. "Many companies report significant improvements." Named, dated, numbered, or delete it. This is the tell that most often indicates a hallucination was smoothed over rather than removed.

The restated conclusion. A final section that repeats the article with the word "ultimately" in it. If the last section contains no information that was not above it, it is padding and the reader can feel it.

Vocabulary that is one notch too formal. Utilize, leverage, delve, robust, seamless. Not wrong, just consistently one register above how anybody talks. The fix is mechanical: search and replace.

No stakes. The deepest tell and the hardest to fix. Real advice includes what happens if you get it wrong, what it costs, what the writer got wrong once. A model has no experience to be burned by. Everything reads as equally safe because to the writer it was.

The sixth is the one to work on. The first five are surface and can be edited out in a pass. The sixth requires somebody who has actually done the thing to add two sentences, and those two sentences are usually the reason the page is worth reading at all.

Where a chat window stops being enough

Everything above works at the scale of one article. The trouble starts on the fifth, and by the fiftieth it is the whole problem.

Voice drifts. Each session starts fresh. You paste your instructions again, slightly differently, and the twelfth article does not sound like the third. Nobody notices from inside — you only see it when a reader reads two pages in a row.

Nothing is reusable. The prompt that worked lives in somebody's scrollback. The improved version lives in a different person's scrollback. There is no shared artefact and no version of it, so the quality of your content depends on which colleague did the writing that week.

Publishing is manual. Copy, paste, reformat, re-add the links that lost their markup, upload the images, set the meta fields. Twenty minutes an article that produces nothing a reader can see. On a weekly cadence that is a day a month of pure transfer.

Nothing connects back to results. The chat does not know which of last quarter's articles gained impressions and which sank, so it cannot tell you what to write next or what to fix. Every piece starts from zero context, which means your content plan never learns.

Old content never gets touched. The archive is where the fastest wins usually are — a page ranking in positions eleven to twenty needs an afternoon, not a new article. But refreshing requires knowing which pages those are, which requires performance data, which the chat window does not have.

Comparison of what breaks at ten, one hundred and one thousand articles when the workflow is a chat window, against what a pipeline has to provide at each of those scales
Nothing on the left is wrong at ten articles. Everything on the left is the whole job at a thousand.

None of this is an argument against using a chat window. It is an argument about where its edge is. Below roughly one article a week with one writer, the tab is genuinely the right tool and adding machinery would be waste. Above that, the constraint stops being the writing and becomes everything around the writing, and no prompt improves any of it. If you are at the point of comparing tooling rather than prompts, the honest survey of AI SEO tools sorts the categories out before you start booking demos.

Four uses that look clever and are not

These come up in every team that gets comfortable with the tool, and each one is a reasonable idea that fails for a reason worth understanding.

Asking it to write the whole article from the keyword. The output is complete, fluent and interchangeable with the output anybody else gets from the same instruction. That is the actual problem: not that it is bad, but that it is average by construction. A model produces the most probable continuation, and the most probable article about a topic is the one that already exists a hundred times. If your differentiator is going to be anything, it has to be information the model did not have — your data, your customers, your mistakes. Give it that and it writes well. Give it a keyword and it writes the average of the internet.

Generating a hundred pages of location or feature variations. Technically easy, and it is the exact behaviour named in the spam documentation: generating many pages without adding value. The test is not whether the pages are different from each other. It is whether each one is worth existing on its own, which for template-filled location pages almost never survives an honest look. There is a legitimate version of programmatic content, and it is legitimate because each page carries real data that differs — inventory, prices, records — rather than the same paragraph with the town name swapped.

Asking it to audit your site. It cannot see your site. Paste a page in and it will comment usefully on that page's text. Ask about crawlability, index coverage, redirect chains or page speed and it is guessing from a URL, which is worth nothing. Those answers live in Search Console, a crawler, and server logs. A model can help you interpret an export you paste in; it cannot generate the export.

Using it as the fact-checker for its own draft. Tempting because it feels like a free second pass, and it does catch real problems: contradictions, missing steps, sections that promise something the article never delivers. What it cannot do is notice that a claim is untrue, because if it had known the claim was untrue it would not have written it. Self-review catches structure. It does not catch facts. Those two things need different reviewers and one of them has to be a person with a source open.

The common thread again: every one of these asks the model for something outside the text it was given. That single test — is this a transformation or a retrieval — sorts almost every question you will ever have about where to draw the line.

A workable weekly routine

Concrete, because abstractions do not survive a busy week. This is a routine for one person responsible for a content programme.

Monday, thirty minutes. Open Search Console. List the queries where you sit between positions eight and twenty with real impressions. That list is your work for the week, and it comes from data, not from a model.

Monday, thirty more. For the top three of those, look at the actual results page. What are the pages that beat you covering? Now paste their visible text into the chat and ask what none of them says. That is a comparison over material you supplied, so it is a safe job.

Tuesday, an hour. Build the brief yourself: the query, the reader, the decision they need to make, the three things you know that the current results do not say. Hand it over for an outline. Argue with the outline. This hour is where the quality of the whole week is decided, and it is the hour people try to skip. The anatomy of a brief that actually constrains a draft is worth reading once and then reusing forever.

Wednesday and Thursday. Draft. Use the model for the parts it does well: rewriting paragraphs that came out clumsy, generating title variants, drafting FAQ questions. Write the parts that need experience yourself — they are usually about a fifth of the text and all of the value.

Friday, an hour. Edit against the six tells. Verify every number by opening its source. Generate the schema, validate it, confirm the page contains what the markup claims. Publish. Then take one old article and improve it, which will frequently outperform the new one.

That routine uses the model for perhaps a third of the total time and none of the judgement. It also scales badly past one person, which is exactly the point of the previous section. The broader habits are covered in the right way to use AI in writing, and the money question — which of these articles is worth the week at all — is the subject of content built to book calls rather than collect traffic.

How this differs from being cited by ChatGPT

Worth separating, because the two get filed under the same heading and they are opposite jobs.

This article is about you using a language model as a tool to do SEO work. The other question — how to get your pages quoted when someone else asks a model about your topic — is a content and structure problem, not a tooling one. It concerns answering questions plainly and early on the page, being consistent about facts across your site, and being the kind of source that gets referenced elsewhere.

They pull in different directions more often than people expect. Optimising to be quotable means putting the plain answer near the top and structuring for extraction. Optimising your workflow with AI means nothing about the page at all. Confusing them produces the worst of both: pages written by a model, structured for a model, and read by nobody.

Where Orova SEO fits, and where it does not

Plainly, since this article has spent two thousand words insisting on plain statements.

Orova SEO is not a chat window and is not a keyword research tool. It does not have a keyword database of its own — you load keywords from a spreadsheet. It holds no backlink index and cannot tell you who links to your competitors. It is not a dedicated rank tracker: positions are read from Google Search Console, so what it knows is what Search Console reports. Its site health check is a scored checklist, not a full technical crawler. If those are the jobs you need, buy the tools that do them.

What it does is the layer this article has been describing as missing. Brand voice set once and stored, so article twelve still sounds like article three. Source material loaded from a Google Drive folder, so the writing is bound to your documents rather than to the model's memory. Drafting on a schedule, then either publishing straight into WordPress or through a publishing API, or holding everything as drafts for a human to read first — the switch flips either way at any time. Refreshing the archive by scanning the posts you already have, finding the weak ones and drafting the optimization outline itself, written over the same URL so the history with Google stays intact. Competitor pages watched, and a technical check-up run against the site.

And the honest limit, which their own FAQ states rather than hides: nobody can promise a date for page one, because Google decides rankings. A publishing system removes the bottleneck of writing and shipping. It does not remove the need to have something worth saying, and it does not substitute for the links and mentions that decide competitive queries.

Common questions

Will Google penalise content written with ChatGPT?

Not for being written with it. The published policy targets content produced for the primary purpose of manipulating rankings, and names automation explicitly while adding that scaled content abuse counts "no matter how it's created". A reviewed, original, genuinely useful article drafted with a model is not the target. A hundred thin pages a week is, and would have been before any of this existed.

Can ChatGPT do keyword research?

It can group and interpret a keyword list you give it, which is genuinely useful. It cannot produce volumes or difficulty scores, and any it offers are invented. Get the numbers from a keyword tool, then hand the list over for clustering and intent hypotheses.

Which model should I use for SEO work?

Less important than the prompt. The gap between a well-constrained prompt and a lazy one is larger than the gap between current frontier models on these tasks. Pick one, build your specification, and change model only when something specific breaks.

Should I disclose that AI helped write an article?

Google's own self-assessment list asks whether the use of automation is self-evident to visitors through disclosures or in other ways, which is a question rather than a rule. The practical answer most teams land on is a real named author who genuinely reviewed the piece and stands behind it. That is what the trust component of E-E-A-T is asking for, and it matters more than a footnote about tooling.

How do I stop it inventing statistics?

Ban them in the prompt, and check anyway. Say explicitly: use no numbers I have not given you, and write "no data" where a figure would go. This reduces the problem and does not eliminate it, because the ban competes with the model's pull toward a complete-looking answer. The only real defence is the rule that nothing goes in the article unless you opened the source yourself.

Is AI-written content as good as human-written content?

Wrong comparison. Almost nothing is written entirely by either, and the interesting variable is where the judgement lives. An article where a person chose the angle, supplied the facts, added what they learned the hard way, and edited the result is a human article that used a tool. An article where the model chose all of that is thin regardless of how it reads.

How much time does this actually save?

On the nine jobs above, a lot — title variants, clustering and rewriting are several times faster. Across a whole article, less than people claim, because the time moves rather than disappearing. You spend less time producing sentences and more time verifying, editing and supplying the parts a model cannot. Teams that report enormous savings have usually stopped doing the verifying.

What to do with this

Take your last three articles and mark every paragraph with who supplied the substance: you, a source you opened, or the model. Then look at how much of the article is in the third bucket.

If it is small, your process is fine and you should spend your effort on the systems problem — voice consistency, publishing, and refreshing what you already have — because that is where your time is going now.

If it is large, no prompt will fix it, and no tool will either. The nine jobs on the safe list all require you to bring something. If you are not bringing anything, the model is not saving you work; it is producing pages that are indistinguishable from every other page produced the same way, and search results are already full of those.

The one habit worth building this week is the smallest one in this article: never publish a number you have not seen at its source. It costs a few minutes per article and it removes the only failure in this whole category that a reader can catch and you cannot.

The chat window is the draft, not the pipeline

Orova SEO is the part that comes after the prompt. Set a brand voice once and every later article still sounds like the earlier ones. Point it at a Google Drive folder so the writing is bound to your own documents. Let it draft, keep everything as drafts for a human to read, or publish straight into WordPress or your own API, with position data read from Search Console.

See Orova SEO