第一次个人编程作业
| 这个作业属于哪个课程 | https://edu.cnblogs.com/campus/gdgy/Networkengineering1834/ |
|---|---|
| 这个作业要求在哪里 | https://edu.cnblogs.com/campus/gdgy/Networkengineering1834/homework/11146 |
| 这个作业的目标 | 实现论文查重,记录PSP表格,使用性能分析器 |
1. GitHub 地址
2. PSP 表格记录程序的各个模块的开发上估计耗费的时间
| PSP2.1 | Personal Software Process Stages | 预估耗时(分钟) |
|---|---|---|
| Planning | 计划 | 60 |
| · Estimate | · 估计这个任务需要多少时间 | 30 |
| Development | 开发 | 400 |
| · Analysis | · 需求分析 (包括学习新技术) | 250 |
| · Design Spec | · 生成设计文档 | 30 |
| · Design Review | · 设计复审 | 20 |
| · Coding Standard | · 代码规范 (为目前的开发制定合适的规范) | 20 |
| · Design | · 具体设计 | 30 |
| · Coding | · 具体编码 | 200 |
| · Code Review | · 代码复审 | 30 |
| · Test | · 测试(自我测试,修改代码,提交修改) | 60 |
| Reporting | 报告 | 40 |
| · Test Repor | · 测试报告 | 40 |
| · Size Measurement | · 计算工作量 | 30 |
| · Postmortem & Process Improvement Plan | · 事后总结, 并提出过程改进计划 | 40 |
| · 合计 | 1280 |
3. 计算模块接口的设计与实现过程
-
1.1 文本转变为字符串类
![]()
读取文件路径,将保留文本汉字,返回字符串 -
1.2 流程类
![]()
输入文件路径,调用 simhash 获取重复率,输出至答案文件路径。 -
1.3 异常类
![]()
用于各种异常的编写。如,若文本为空,则抛出异常。
- 1.4 simhash 算法类
![]()
simhash算法的输入是一个向量,输出是一个 f 位的签名值。为了陈述方便,假设输入的是一个文档的特征集合,每个特征有一定的权重。比如特征可以是文档中的词,其权重可以是这个词出现的次数。
simhash 算法如下:
1,将一个 f 维的向量 V 初始化为 0 ; f 位的二进制数 S 初始化为 0 ;
2,对每一个特征:用传统的 hash 算法对该特征产生一个 f 位的签名 b 。对 i=1 到 f :
如果b 的第 i 位为 1 ,则 V 的第 i 个元素加上该特征的权重;
否则,V 的第 i 个元素减去该特征的权重。
3,如果 V 的第 i 个元素大于 0 ,则 S 的第 i 位为 1 ,否则为 0 ;
4,输出 S 作为签名。

比较相似度:
海明距离: 两个码字的对应比特取值不同的比特数称为这两个码字的海明距离。一个有效编码集中, 任意两个码字的海明距离的最小值称为该编码集的海明距离。
举例如下: 10101 和 00110 从第一位开始依次有第一位、第四、第五位不同,则海明距离为 3.
对每篇文档根据SimHash 算出签名后,再计算两个签名的海明距离(两个二进制异或后 1 的个数)即可。
根据经验值,对 64 位的 SimHash ,海明距离在 3 以内的可以认为相似度比较高。
1.5 程序流程图

1.6 类与方法的调用

4. 计算模块接口部分的性能改进
-
Overview
![]()
-
Live Memory
![]()
hanlp 和 大数 占用内存较多
消耗最大的函数是 SimHash
5. 计算模块部分单元测试展示
- 测试代码
package org.example;
import org.junit.Test;
import java.io.File;
import java.io.FileWriter;
public class test
{
@Test
public void addTest() {
String []s = {"text\\orig.txt", "text\\orig_0.8_add.txt", "text\\ans.txt"};
try {
File file = new File(s[2]);
if (!file.exists()) {
boolean flag = file.createNewFile();
if (!flag) throw new Exception();
}
FileWriter fileWriter = new FileWriter(file.getAbsoluteFile(), true);
fileWriter.write("与orig_0.8_add.txt重复率: ");
fileWriter.close();
} catch (Exception e) {
e.printStackTrace();
}
Main.main(s);
}
@Test
public void delTest() {
String []s = {"text\\orig.txt", "text\\orig_0.8_del.txt", "text\\ans.txt"};
try {
File file = new File(s[2]);
if (!file.exists()) {
boolean flag = file.createNewFile();
if (!flag) throw new Exception();
}
FileWriter fileWriter = new FileWriter(file.getAbsoluteFile(), true);
fileWriter.write("与orig_0.8_del.txt重复率: ");
fileWriter.close();
} catch (Exception e) {
e.printStackTrace();
}
Main.main(s);
}
@Test
public void dis_1Test() {
String []s = {"text\\orig.txt", "text\\orig_0.8_dis_1.txt", "text\\ans.txt"};
try {
File file = new File(s[2]);
if (!file.exists()) {
boolean flag = file.createNewFile();
if (!flag) throw new Exception();
}
FileWriter fileWriter = new FileWriter(file.getAbsoluteFile(), true);
fileWriter.write("与orig_0.8_dis_1.txt重复率: ");
fileWriter.close();
} catch (Exception e) {
e.printStackTrace();
}
Main.main(s);
}
@Test
public void dis_10Test() {
String []s = {"text\\orig.txt", "text\\orig_0.8_dis_10.txt", "text\\ans.txt"};
try {
File file = new File(s[2]);
if (!file.exists()) {
boolean flag = file.createNewFile();
if (!flag) throw new Exception();
}
FileWriter fileWriter = new FileWriter(file.getAbsoluteFile(), true);
fileWriter.write("与orig_0.8_dis_10.txt重复率: ");
fileWriter.close();
} catch (Exception e) {
e.printStackTrace();
}
Main.main(s);
}
@Test
public void dis_15Test() {
String []s = {"text\\orig.txt", "text\\orig_0.8_dis_15.txt", "text\\ans.txt"};
try {
File file = new File(s[2]);
if (!file.exists()) {
boolean flag = file.createNewFile();
if (!flag) throw new Exception();
}
FileWriter fileWriter = new FileWriter(file.getAbsoluteFile(), true);
fileWriter.write("与orig_0.8_dis_15.txt重复率: ");
fileWriter.close();
} catch (Exception e) {
e.printStackTrace();
}
Main.main(s);
}
@Test
public void sameTest() {
String []s = {"text\\orig.txt", "text\\orig.txt", "text\\ans.txt"};
try {
File file = new File(s[2]);
if (!file.exists()) {
boolean flag = file.createNewFile();
if (!flag) throw new Exception();
}
FileWriter fileWriter = new FileWriter(file.getAbsoluteFile(), true);
fileWriter.write("与orig.txt的重复率: ");
fileWriter.close();
} catch (Exception e) {
e.printStackTrace();
}
Main.main(s);
}
@Test
public void emptyTest() {
String s = "text\\orig.txt";
String t = "text\\empty.txt";
String ansPath = "text\\ans.txt";
Solve.solve(s, t, ansPath);
}
@Test
public void filePathTest() {
String s = "text\\orig.tx";
String t = "text\\o.txt";
String ansPath = "text\\ans.txt";
Solve.solve(s, t, ansPath);
}
}
- 测试结果
测试全部通过


- 测试覆盖率:
![]()

6. 计算模块部分异常处理说明
- 设计异常信息传递类
public class MyException extends Exception {
public MyException(String s) {
super(s);
}
}
- 测试样例
-
空文本
![]()
-
文件路径错误
![]()
-
7. PSP表格记录下你在程序的各个模块上实际花费的时间
| PSP2.1 | Personal Software Process Stages | 预估耗时(分钟) | 实际耗时(分钟) |
|---|---|---|---|
| Planning | 计划 | 60 | 50 |
| · Estimate | · 估计这个任务需要多少时间 | 30 | 25 |
| Development | 开发 | 400 | 500 |
| · Analysis | · 需求分析 (包括学习新技术) | 250 | 300 |
| · Design Spec | · 生成设计文档 | 30 | 50 |
| · Design Review | · 设计复审 | 20 | 10 |
| · Coding Standard | · 代码规范 (为目前的开发制定合适的规范) | 20 | 10 |
| · Design | · 具体设计 | 30 | 40 |
| · Coding | · 具体编码 | 200 | 240 |
| · Code Review | · 代码复审 | 30 | 20 |
| · Test | · 测试(自我测试,修改代码,提交修改) | 60 | 60 |
| Reporting | 报告 | 40 | 30 |
| · Test Repor | · 测试报告 | 40 | 40 |
| · Size Measurement | · 计算工作量 | 30 | 25 |
| · Postmortem & Process Improvement Plan | · 事后总结, 并提出过程改进计划 | 40 | 50 |
| · 合计 | 1280 | 1450 |
8. 总结
从本次的个人项目中收获了许多知识:寻找算法, 编写代码,使用性能分析器 JProfile, 用 Junit 进行单元测试。
面向百度编程,查找所遇到的问题的解决方法,不断学习新东西,一步步完善项目。










浙公网安备 33010602011771号