I recently asked an AI to evaluate how well I communicate with it. The answer was detailed, useful, and mostly fair. It scored different parts of my prompting style, including how I give context, define goals, correct misunderstandings, delegate decisions, separate ideas from commitments, and establish acceptance criteria.
Some of the scores were high. Others were much lower.
At first, the lower scores seemed reasonable. I often give broad instructions. I move between ideas quickly. I do not always define exactly what the finished result should look like before beginning. I may introduce several possibilities in the same conversation, change direction, or expect the AI to fill in technical blanks.
From the outside, that can look like weak prompting.
But once I explained why I work that way, some of the scores changed dramatically.
The most obvious example was decision authority. I had initially been scored around 5 out of 10 because I do not always specify which decisions the AI is allowed to make. The assumption was that I was leaving important choices unclear.
That was not what I was doing.
When I say “fill in the blanks,” I am intentionally delegating. I may not know the correct architecture, terminology, mechanism, or implementation path. That is one of the reasons I am asking an AI research assistant in the first place. I want it to use its broader knowledge, consider a spectrum of logical possibilities, account for what it already knows about me and my projects, and make a real recommendation.
I do not need to own every technical decision. I need to understand the result well enough to recognize whether it fits, identify what is wrong, and steer it when necessary.
Once I explained that, the score rose from roughly 5 out of 10 to 9 out of 10.
The behavior had not changed. The interpretation had.
The same thing happened with my tendency to mix exploration and commitment. I often think through many ideas at once. I may have several conversations open, multiple projects moving, several tools in use, and many possible future paths competing for attention.
Initially, that was judged as poor separation between brainstorming and implementation.
There was truth in the observation, but the explanation mattered. This is not simply disorganization. It is part of how I design systems. I think across layers, dependencies, future consequences, interfaces, risks, and alternate possibilities at the same time.
Trying to force that process into a perfectly linear sequence would make it easier to document, but it could also reduce the creativity that makes it valuable.
The better conclusion was not that I needed to stop thinking that way. It was that the AI needed to become better at organizing the flow.
My role is to produce ideas, connections, concerns, and direction. Its role is to help distinguish what belongs in the current build, what belongs on the roadmap, what is speculative, what has been decided, and what should be rejected.
Again, the behavior did not change. The model of the behavior changed.
Acceptance criteria created another disagreement. I had been scored very low because I do not always define in advance how a project will be judged complete.
My response was simple: sometimes I do not know yet.
When a concept is still being discovered, rigid acceptance criteria can become premature constraints. They can force an emerging idea to fit the limits of what is already understood.
That does not mean testing is unimportant. It means the tests may need to emerge with the design.
During exploration, the purpose may be to find a coherent direction worth pursuing. During architecture, the important truths may be boundaries, responsibilities, data flows, and failure behavior. During implementation, those discoveries can finally become concrete tests.
The AI can help derive those tests. I should not always have to provide them before the design exists.
The lowest scores improved because I provided context that was not visible in the original behavior.
That revealed a larger lesson about evaluating communication.
A prompt cannot always be judged only by how formally complete it appears. The same sentence can mean different things depending on the relationship, the history of the project, the user’s knowledge, and the amount of authority being intentionally delegated.
“Figure it out” can be careless.
It can also mean:
Use your expertise. Make the strongest reasonable decision. Explain your reasoning. Challenge me where necessary. I will evaluate the result and steer the next step.
Those are very different instructions, even if they initially look similar.
The reassessment also showed that effective human-AI collaboration is not only about writing perfect standalone prompts. It can also be built through trust, accumulated context, correction, and division of cognitive labor.
I do not expect the AI merely to execute decisions I have already made. I expect it to contribute informed opinions, identify weaknesses, make connections I missed, and push back when my assumptions are poor.
I expect it to teach me while we work.
In return, I provide direction, context, intuition, correction, and an understanding of the larger system I am trying to build.
The process is not perfectly linear. It is iterative. Sometimes I begin with a clear requirement. Sometimes I begin with an incomplete thought. Sometimes the idea only becomes understandable after the first attempt is wrong.
That does not necessarily indicate failure.
Sometimes the first wrong answer is what gives the idea enough shape to be corrected.
The most important outcome of the discussion was not that I received higher scores. It was that both sides developed a more accurate understanding of the working relationship.
The original evaluation measured my behavior against a conventional model of prompt writing.
The revised evaluation measured it against what I was actually trying to accomplish.
That difference matters.
A low score may indicate a real weakness. But it may also indicate that the evaluator has not yet understood the system it is evaluating.
Sometimes better prompting means giving clearer instructions.
Sometimes it means explaining why the unusual instructions are intentional.

