学习selenium

学习selenium

参照:https://blog.csdn.net/IT_LanTian/article/details/122986725

非常nb的博主

下载

pip install selenium

下载驱动

(1)Chrome浏览器驱动(chromedriver ):http://chromedriver.storage.googleapis.com/index.html

(2)Firefox浏览器驱动(geckodriver):https://github.com/mozilla/geckodriver/releases/

(3)Edge浏览器驱动(MicrosoftWebDriver):https://github.com/mozilla/geckodriver/releases/

(4)IE浏览器驱动(IEDriverServer):http://selenium-release.storage.googleapis.com/index.html

(5)Opera浏览器驱动(operadriver):https://github.com/operasoftware/operachromiumdriver/releases

(6)PhantomJS浏览器驱动(phantomjs):https://github.com/operasoftware/operachromiumdriver/releases

将驱动放在python解释器的根目录(python解释器在哪一个位置,放在同目录下),才能使用驱动

学习

之前学过selenium的,很多方法都会过时

比如:

属性定位方法 原定位方法find_element_by_ 推荐定位方法find_element()
xpath find_element_by_xpath(“//*[@id=‘search’]”) find_element(By.XPATH, “//*[@id=‘search’]”)
class_name find_element_by_class_name(“element_class_name”) find_element(By.CLASS_NAME, “element_class_name”)
id find_element_by_id(“element_id”) find_element(By.ID,“element_id”)
name find_element_by_name(“element_name”) find_element(By.NAME, “element_name”)
link_text find_element_by_link_text(“element_link_text”) find_element(By.LINK_TEXT,“element_link_text”)
css_selector find_element_by_css_selector(“element_css_selector”) find_element(By.CSS_SELECTOR, “element_css_selector”)
tag_name find_element_by_tag_name(“element_tag_name”) find_element(By.TAG_NAME, “element_tag_name”)
partial_link_text find_element_by_partial_link_text(“element_partial_link_text”) find_element(By.PARTIAL_LINK_TEXT, “element_partial_link_text”)

上面的方法一定要记住,这是自动化查找页面元素的基础,尤其是第三列

本博客使用的selenium是4.16版本,当前最新版

开始学习

导入库

# 进行自动化的关键库
from selenium import webdriver

# 新型查找元素的库
from selenium.webdriver.common.by import By

# 一些需要键盘来操作界面的keys,最常用的是enter
# 比如 search.send_keys(Keys.ENTER)
from selenium.webdriver.common.keys import Keys

加载驱动

以下是最常用的浏览器驱动

# edge ie升级版本,非常好用
# 可以在webdriver.Edge()函数中放入驱动的绝对路径
edge = webdriver.Edge()
# chrome 谷歌 最好用的浏览器
chrome = webdriver.Chrome()
# firefox 火狐 比较好用的浏览器
firefox = webdriver.Firefox()
#  safari苹果设备自带浏览器
safari = webdriver.Safari()

获取网页,改变网页状态

# 会在电脑上打开你输入的网址
chrome.get("https://www.baidu.com")

# 使网页达到全屏
# 否者一些元素被隐藏,导致页面元素不完整
chrome.maximize_window()


# 设置分辨率 500*500
browser.set_window_size(500,500)

获取网页元素--最重要的一环

# find_element()函数 相比于旧版更好用

#以下都可以为查找页面元素做贡献,但最常用的还是XPATH
#By.ID
#By.XPATH  最常使用的,页面有唯一属性
#By.CLASS_NAME
#By.NAME
#BY.CSS_SELECTOR   这种方法相对xpath要简洁些,定位速度也要快些 #kw
#BY.LINK_TEXT  用来定位文本链接
#By.PARTIAL_LINK_TEXT 有时候一个超链接的文本很长,我们如果全部输入,既麻烦,又显得代码很不美观,这时候我们就可以只截取一部分字符串,用这种方法模糊匹配了
#By.TAG_NAME  范围太大,经常出错,避免使用

chrome.find_element(By.ID, "kw").send_keys("java")
chrome.find_element(By.ID, "su").click()

# 如果定位的目标元素在网页中不止一个,那么则需要用到find_elements,得到的结果会是列表形式
chrome.find_elements()

动作

# 最常用的几种动作

# 输入
chrome.find_element(By.ID, "kw").send_keys("java")
# 点击
chrome.find_element(By.ID, "kw").click()
# 在搜索框输入文本python,然后回车就出查询操作结果的情况
chrome.find_element(By.ID, "kw").submit()
# 既然有输入,这里也就有清除文本啦
chrome.find_element(By.ID, "kw").clear()

# 判断是否显示
chrome.find_element(By.ID, "kw").is_displayed()
# 判断是否enable
chrome.find_element(By.ID, "kw").is_enabled()
#判断是否选择
chrome.find_element(By.ID, "kw").is_selected()

退出

# Closes the browser and shuts down the ChromiumDriver executable
# 关闭浏览器
chrome.quit()

# Closes the current window
# 关闭当前页面
edge.close()

拍照

# 测试的时候需要保留失败的照片
# 拍照的是当前页面
# 下面两个方法,用法一摸一样
edge.save_screenshot("照片.jpg")

edge.get_screenshot_as_file("照片.png")

等待

# 有时候页面加载慢,会导致信息不匹配,需要等待
# 很重要!
time.sleep(3)

选项卡跳转

# 有时候需要跳转到以前的页面
# 以前的页面会存在一个列表中

edge.window_handles
# 第一个字符串相当于第一个页面的唯一id号
#['8801DEE42FD974B0640FBAA34C5A2918', '9938F835EA06E33EEBCF0A5F074BCCD6']

edge.current_window_handle
# 当前页面的唯一id号
#'8801DEE42FD974B0640FBAA34C5A2918'

edge.switch_to.window(edge.window_handles[0])
# 根据id号选择跳转到哪一个页面

#跳转时有可能句柄还为加载,最好在前面加一个sleep()等待

前进,后退

与箭头左,箭头右作用相同

browser.get(r'https://www.baidu.com')  
time.sleep(2)
 
# 打开淘宝页面
browser.get(r'https://www.taobao.com')  
time.sleep(2)
 
# 后退到百度页面
browser.back()  
time.sleep(2)
 
# 前进的淘宝页面
browser.forward() 
time.sleep(2)

一些重要的页面属性

edge = webdriver.Edge()
edge.get("https://www.baidu.com")
print(edge.title)# 百度一下,你就知道
print(edge.page_source[0:10])# <html><hea
# 页面源码我们就可以用正则表达式、Bs4、xpath以及pyquery等工具进行解析提取想要的信息了

print(edge.name)# msedge
print(edge.current_url)# https://www.baidu.com/

在测试场景中

有一种场景是在浏览器的各个版本中进行测试

那么就需要浏览器各个版本的驱动

一个一个下载非常麻烦

有一个自动下载驱动的库非常方便

千万要注意:selenium版本很重要,会影响下载

下面的程序是在4.7版本运行

4.16版本会有很多bug导致无法运行,请自行选择版本

pip install webdriver-manager
from webdriver_manager.chrome import ChromeDriverManager

c=ChromeDriverManager()
# 会自动查看浏览器版本,下载相应的驱动
s=c.install()

# 下载的地址会放到
# C:\Users\14247\.wdm\drivers\chromedriver\win64\119.0.6045.105\chromedriver-win32/chromedriver.exe
print(s)

browser = webdriver.Chrome(s)

# 火狐比较特殊,edge与上述一样
# browser = webdriver.Chrome(executable_path=path)

如果因为网速问题下载过慢,可以使用镜像

进入ChromeDriverManager

将两行修改地址

url: str = "https://chromedriver.storage.googleapis.com",
            latest_release_url: str = "https://chromedriver.storage.googleapis.com/LATEST_RELEASE",

为

url="http://npm.taobao.org/mirrors/chromedriver",
latest_release_url="http://npm.taobao.org/mirrors/chromedriver/LATEST_RELEASE",

优化

有时候我们不需要页面展示

可以使用无界面化

# 创建参数对象
options = webdriver.ChromeOptions()
# 添加参数
options.add_argument("headless")
# 为驱动添加参数
browser = webdriver.Chrome(s,options=options)

有时候需要获取新的数据

需要刷新页面

try:
    # 刷新页面
    browser.refresh()  
    print('刷新页面')
except Exception as e:
    print('刷新失败')

有时候获取的网页有很多有价值的内容

可以爬取他们

browser = webdriver.Chrome()
 
browser.get(r'https://www.baidu.com')  
 
logo = browser.find_element_by_class_name('index-logo-src')
print(logo)
print(logo.get_attribute('src'))

#<selenium.webdriver.remote.webelement.WebElement (session="e95b18c43a330745af019e0041f0a8a4", element="7dad5fc0-610b-45b6-b543-9e725ee6cc5d")>
#https://www.baidu.com/img/PCtm_d9c8750bed0b3c7d089fa7d55720d6cf.png

logo = browser.find_element_by_css_selector('#hotsearch-content-wrapper > li:nth-child(1) > a')
print(logo.text)
print(logo.get_attribute('href'))

#1各地贯彻十九届六中全会精神纪实
#https://www.baidu.com/s?cl=3&tn=baidutop10&fr=top1000&wd=%E5%90%84%E5%9C%B0%E8%B4%AF%E5%BD%BB%E5%8D%81%E4%B9%9D%E5%B1%8A%E5%85%AD%E4%B8%AD%E5%85%A8%E4%BC%9A%E7%B2%BE%E7%A5%9E%E7%BA%AA%E5%AE%9E&rsv_idx=2&rsv_dl=fyb_n_homepage&sa=fyb_n_homepage&hisfilter=1

logo = browser.find_element_by_class_name('index-logo-src')
print(logo.id)
print(logo.location)
print(logo.tag_name)
print(logo.size)

# 6af39c9b-70e8-4033-8a74-7201ae09d540
# {'x': 490, 'y': 46}
# img
# {'height': 129, 'width': 270}

下拉框

下拉框比较特殊

需要单独使用一章来进行介绍

需要导入

from selenium.webdriver.support.select import Select

自制页面进行学习

<!DOCTYPE html>
<html lang="en">
<head>
    <meta charset="UTF-8">
    <title>Title</title>
</head>
<body>
<form>
<select name="fruit">
<option value="apple">苹果</option>
<option value="banana" selected="">香蕉</option>
<option value="watermelon">西瓜</option>
<option value="grape">葡萄</option>
</select>
</form>

</body>
</html>

基础知识点

'''1、三种选择某一选项项的方法'''
 
select_by_index()           # 通过索引定位;注意:index索引是从“0”开始。
select_by_value()           # 通过value值定位,value标签的属性值。
select_by_visible_text()    # 通过文本值定位,即显示在下拉框的值。
 
'''2、三种返回options信息的方法'''
# 有问题
# 返回的是<selenium.webdriver.remote.webelement.WebElement (session="d2a670c378f79d96f6ee8d1d1f001f3b", element="7A3518517C498D98F256AF4BD45EC245_element_6")>
 
options                     # 返回select元素所有的options
all_selected_options        # 返回select元素中所有已选中的选项
first_selected_options      # 返回select元素中选中的第一个选项                  
 
 
'''3、四种取消选中项的方法''' 
 # 以下全是对多选框有效
deselect_all                # 取消全部的已选择项 
deselect_by_index           # 取消已选中的索引项
deselect_by_value           # 取消已选中的value值
deselect_by_visible_text    # 取消已选中的文本值

实操

browser = webdriver.Chrome(s)

browser.get(r'E:\pythonstudy\project\testcase\x.html')

Select(browser.find_element(By.NAME, "fruit")).select_by_index(0)

time.sleep(2)
Select(browser.find_element(By.NAME, "fruit")).select_by_index(1)

time.sleep(2)
Select(browser.find_element(By.NAME, "fruit")).select_by_index(2)

time.sleep(2)
Select(browser.find_element(By.NAME, "fruit")).select_by_index(3)

time.sleep(2)
# 关闭浏览器
browser.close()

多窗口切换

比如同一个页面的不同子页面的节点元素获取操作,不同选项卡之间的切换以及不同浏览器窗口之间的切换操作等等。

Frame切换

Selenium打开一个页面之后,默认是在父页面进行操作,此时如果这个页面还有子页面,想要获取子页面的节点元素信息则需要切换到子页面进行擦走,这时候switch_to.frame()就来了。如果想回到父页面,用switch_to.parent_frame()即可。

选项卡切换

我们在访问网页的时候会打开很多个页面,在Selenium中提供了一些方法方便我们对这些页面进行操作。
current_window_handle:获取当前窗口的句柄。

window_handles:返回当前浏览器的所有窗口的句柄。

switch_to_window():用于切换到对应的窗口。

案例

from selenium import webdriver
import time
 
browser = webdriver.Chrome()
 
# 打开百度
browser.get('http://www.baidu.com')
# 新建一个选项卡
browser.execute_script('window.open()')
print(browser.window_handles)
# 跳转到第二个选项卡并打开知乎
browser.switch_to.window(browser.window_handles[1])
browser.get('http://www.zhihu.com')
# 回到第一个选项卡并打开淘宝(原来的百度页面改为了淘宝)
time.sleep(2)
browser.switch_to.window(browser.window_handles[0])
browser.get('http://www.taobao.com')

模拟鼠标操作

导入库

from selenium.webdriver.common.action_chains import ActionChains

左键单击

click()

右键单击

context_click()
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.by import By
from selenium import webdriver
import time

edge = webdriver.Edge()
edge.get(r"https://www.baidu.com")
time.sleep(2)

right_click = edge.find_element(By.LINK_TEXT, "新闻")
ActionChains(edge).context_click(right_click).perform()

time.sleep(2)
edge.close()

函数说明

ActionChains(browser):调用ActionChains()类,并将浏览器驱动browser作为参数传入

context_click(right_click):模拟鼠标双击,需要传入指定元素定位作为参数

perform():执行ActionChains()中储存的所有操作,可以看做是执行之前一系列的操作

双击左键

double_click()
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.by import By
from selenium import webdriver
import time

edge = webdriver.Edge()
edge.get(r"https://www.baidu.com")
time.sleep(2)

double_click = edge.find_element(By.XPATH,'//*[@id="hotsearch-refresh-btn"]/span')
ActionChains(edge).double_click(double_click).perform()

time.sleep(2)
edge.close()

拖拽

drag_and_drop(source,target)拖拽操作嘛,开始位置和结束位置需要被指定,这个常用于滑块类验证码的操作之类
from selenium.webdriver.common.action_chains import ActionChains
from selenium import webdriver
import time  
 
browser = webdriver.Chrome()
url = 'https://www.runoob.com/try/try.php?filename=jqueryui-api-droppable'
browser.get(url)  
time.sleep(2)
 
browser.switch_to.frame('iframeResult')
 
# 开始位置
source = browser.find_element_by_css_selector("#draggable")
 
# 结束位置
target = browser.find_element_by_css_selector("#droppable")
 
# 执行元素的拖放操作
actions = ActionChains(browser)
actions.drag_and_drop(source, target)
actions.perform()
# 拖拽
time.sleep(15)
 
# 关闭浏览器
browser.close()

悬停

move_to_element()
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.by import By
from selenium import webdriver
import time

edge = webdriver.Edge()
edge.get(r"https://www.baidu.com")
time.sleep(2)

move = edge.find_element(By.XPATH,'//*[@id="s-hotsearch-wrapper"]/div/a[1]/div/i[1]')
ActionChains(edge).move_to_element(move).perform()

time.sleep(4)
edge.close()

模拟键盘

引入类

from selenium.webdriver.common.keys import Keys

常见键盘操作

send_keys(Keys.BACK_SPACE):删除键(BackSpace)

send_keys(Keys.SPACE):空格键(Space)

send_keys(Keys.TAB):制表键(TAB)

send_keys(Keys.ESCAPE):回退键(ESCAPE)

send_keys(Keys.ENTER):回车键(ENTER)

send_keys(Keys.CONTRL,'a'):全选(Ctrl+A)

send_keys(Keys.CONTRL,'c'):复制(Ctrl+C)

send_keys(Keys.CONTRL,'x'):剪切(Ctrl+X)

send_keys(Keys.CONTRL,'v'):粘贴(Ctrl+V)

send_keys(Keys.F1):键盘F1

.....

send_keys(Keys.F12):键盘F12

定位需要操作的元素,然后操作即可!

from selenium.webdriver.common.keys import Keys
from selenium import webdriver
import time
 
browser = webdriver.Chrome()
url = 'https://www.baidu.com'
browser.get(url)  
time.sleep(2)
 
# 定位搜索框
input = browser.find_element_by_class_name('s_ipt')
# 输入python
input.send_keys('python')
time.sleep(2)
 
# 回车
input.send_keys(Keys.ENTER)
time.sleep(5)
 
# 关闭浏览器
browser.close()

延迟等待

以下是几种等待方式

强制等待

time.sleep()

隐式等待

implicitly_wait()设置等待时间,如果到时间有元素节点没有加载出来,就会抛出异常。
from selenium import webdriver
 
browser = webdriver.Chrome()
 
# 隐式等待,等待时间10秒
browser.implicitly_wait(10)  
 
browser.get('https://www.baidu.com')
print(browser.current_url)
print(browser.title)
 
# 关闭浏览器
browser.close()

显式等待

最常用,能使效率最大化

设置一个等待时间和一个条件,在规定时间内,每隔一段时间查看下条件是否成立,如果成立那么程序就继续执行,否则就抛出一个超时异常

导入两个类

# 设置等待时间
from selenium.webdriver.support.wait import WebDriverWait
# 设置等待时间所需要的条件
from selenium.webdriver.support import expected_conditions as EC

示例

from selenium import webdriver
from selenium.webdriver.support.wait import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
import time
 
browser = webdriver.Chrome()
browser.get('https://www.baidu.com')
# 设置等待时间10s
wait = WebDriverWait(browser, 10)
# 设置判断条件:等待id='kw'的元素加载完成
input = wait.until(EC.presence_of_element_located((By.ID, 'kw')))
# 在关键词输入:关键词
input.send_keys('Python')
 
# 关闭浏览器
time.sleep(2)
browser.close()

参数说明

WebDriverWait(driver,timeout,poll_frequency=0.5,ignored_exceptions=None)

driver: 浏览器驱动

timeout: 超时时间,等待的最长时间(同时要考虑隐性等待时间)

poll_frequency: 每次检测的间隔时间,默认是0.5秒

ignored_exceptions:超时后的异常信息,默认情况下抛出NoSuchElementException异常

until(method,message='')

method: 在等待期间,每隔一段时间调用这个传入的方法,直到返回值不是False

message: 如果超时,抛出TimeoutException,将message传入异常

until_not(method,message='')

until_not 与until相反,until是当某元素出现或什么条件成立则继续执行,until_not是当某元素消失或什么条件不成立则继续执行,参数也相同。

其他等待条件

from selenium.webdriver.support import expected_conditions as EC
 
# 判断标题是否和预期的一致
title_is
# 判断标题中是否包含预期的字符串
title_contains
 
# 判断指定元素是否加载出来
presence_of_element_located
# 判断所有元素是否加载完成
presence_of_all_elements_located
 
# 判断某个元素是否可见. 可见代表元素非隐藏,并且元素的宽和高都不等于0,传入参数是元组类型的locator
visibility_of_element_located
# 判断元素是否可见,传入参数是定位后的元素WebElement
visibility_of
# 判断某个元素是否不可见,或是否不存在于DOM树
invisibility_of_element_located
 
# 判断元素的 text 是否包含预期字符串
text_to_be_present_in_element
# 判断元素的 value 是否包含预期字符串
text_to_be_present_in_element_value
 
#判断frame是否可切入,可传入locator元组或者直接传入定位方式:id、name、index或WebElement
frame_to_be_available_and_switch_to_it
 
#判断是否有alert出现
alert_is_present
 
#判断元素是否可点击
element_to_be_clickable
 
# 判断元素是否被选中,一般用在下拉列表,传入WebElement对象
element_to_be_selected
# 判断元素是否被选中
element_located_to_be_selected
# 判断元素的选中状态是否和预期一致,传入参数:定位后的元素,相等返回True,否则返回False
element_selection_state_to_be
# 判断元素的选中状态是否和预期一致,传入参数:元素的定位,相等返回True,否则返回False
element_located_selection_state_to_be
 
#判断一个元素是否仍在DOM中,传入WebElement对象,可以判断页面是否刷新了
staleness_of

使用脚本

还有一些操作,比如下拉进度条,模拟javaScript,使用execute_script方法来实现
from selenium import webdriver
 
browser = webdriver.Chrome()
# 知乎发现页
browser.get('https://www.zhihu.com/explore')
 
browser.execute_script('window.scrollTo(0, document.body.scrollHeight)')
browser.execute_script('alert("To Bottom")')

cookie的获取

cookie在爬虫中非常有用

# 获取所有cookie
browser.get_cookies()
# 添加cookie
browser.add_cookie()
# 删除cookie
browser.delete_cookie()
# 删除所有cookie
browser.delete_all_cookies()
posted @ 2023-12-14 11:22  Bre-eZe  阅读(89)  评论(0)    收藏  举报